Claude Skill

ai-provider-mistral-sdk

Official Mistral AI TypeScript SDK patterns — client setup, chat completions, streaming, function calling, structured outputs, embeddings, vision, Codestral FIM, and production best practices

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_ai-provider-mistral-sdk_skills_ai-provider-mistral-sdk-3a51ef5.zip · 22 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-provider-mistral-sdk/skills/ai-provider-mistral-sdk
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Mistral SDK Patterns

Quick Guide: Use @mistralai/mistralai (ESM-only) to interact with Mistral's API. Use client.chat.complete() for chat, client.chat.stream() for streaming (async iterable via for await), client.chat.parse() with a Zod schema for structured outputs, and client.fim.complete() for Codestral fill-in-middle code completion. The SDK uses responseFormat (camelCase) not response_format. Streaming events expose content via event.data.choices[0]?.delta?.content. Retries default to strategy: "none" -- you must configure them explicitly for production.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST use responseFormat (camelCase) in SDK calls -- NOT response_format (snake_case). The SDK uses camelCase property names throughout.)

(You MUST configure retries explicitly -- the SDK defaults to strategy: "none" (no retries), unlike OpenAI's SDK which retries automatically)

(You MUST consume streaming results with for await (const event of result) and access content via event.data.choices[0]?.delta?.content -- the event shape differs from OpenAI)

(You MUST never hardcode API keys -- use process.env["MISTRAL_API_KEY"] with the bracket notation the SDK documents)

(You MUST use client.chat.parse() with a Zod schema for structured outputs -- NOT manual JSON.parse() on completion content)

</critical_requirements>


Auto-detection: Mistral, mistral, @mistralai/mistralai, client.chat.complete, client.chat.stream, client.chat.parse, client.fim.complete, client.embeddings.create, mistral-large, mistral-small, codestral, pixtral, ministral, magistral, devstral, MISTRAL_API_KEY, responseFormat, mistral-embed

When to use:

  • Building applications that call Mistral models directly (Mistral Large, Small, Codestral, etc.)
  • Implementing chat completions with SSE streaming
  • Using Codestral for code generation and fill-in-middle (FIM) completion
  • Extracting structured data with client.chat.parse() and Zod schemas
  • Implementing function calling / tool use
  • Creating embeddings for RAG pipelines or semantic search
  • Processing images with vision-capable models (Mistral Small, Medium, Large, Ministral)
  • Using Mistral Agents API for pre-configured agent completions

Key patterns covered:

  • Client initialization and configuration (retries, timeouts, custom HTTP client)
  • Chat completions (chat.complete) and streaming (chat.stream)
  • Structured outputs with chat.parse() and Zod schemas
  • Function calling / tool use with tool call loop
  • Embeddings (embeddings.create) with mistral-embed
  • Vision (image URL / base64 with vision-capable models)
  • Codestral FIM (fim.complete) for code completion
  • Error handling, retry configuration, and production patterns

When NOT to use:

  • Multi-provider applications where you need to switch between Mistral, OpenAI, Anthropic, etc. -- use a unified provider SDK
  • React-specific chat UI hooks (useChat) -- use a framework-integrated AI SDK
  • When you need OpenAI-compatible endpoints -- use OpenAI SDK with Mistral's compatible endpoint instead

Examples Index





<decision_framework>

Decision Framework

Which Method to Use

What do you need?
+-- Chat completion (text in, text out)?
|   +-- Need streaming? -> client.chat.stream()
|   +-- Need structured JSON? -> client.chat.parse() with Zod schema
|   +-- Basic completion? -> client.chat.complete()
+-- Code completion / fill-in-middle?
|   +-- YES -> client.fim.complete() with Codestral
+-- Embeddings for search/RAG?
|   +-- YES -> client.embeddings.create() with mistral-embed
+-- Pre-configured agent?
    +-- YES -> client.agents.complete() with agent ID

Which Model to Choose

What is your task?
+-- Most capable general purpose -> mistral-large-latest
+-- Balanced cost/performance -> mistral-medium-latest
+-- Fast + cost-efficient -> mistral-small-latest
+-- Minimal / edge deployment -> ministral-3b-latest
+-- Complex reasoning / math -> magistral-medium-latest
+-- Code generation (chat) -> codestral-latest or devstral-latest
+-- Code completion (FIM) -> codestral-latest
+-- Vision / image analysis -> mistral-small-latest (or any vision-capable model)
+-- Embeddings -> mistral-embed
+-- Code embeddings -> codestral-embed-latest

Streaming vs Non-Streaming

Is the response user-facing?
+-- YES -> Use client.chat.stream()
|   +-- Iterate with: for await (const event of result)
|   +-- Access content: event.data.choices[0]?.delta?.content
+-- NO -> Use client.chat.complete()
    +-- Background processing -> chat.complete()
    +-- Structured output -> chat.parse() with Zod

When to Use This SDK vs a Provider-Agnostic SDK

Do you need multiple LLM providers (Mistral + others)?
+-- YES -> Not this skill's scope -- use a unified provider SDK
+-- NO -> Do you need Mistral-specific features?
    +-- YES -> Use Mistral SDK directly
    |   Examples: Codestral FIM, Mistral Agents,
    |   Voxtral audio, OCR, custom endpoints
    +-- NO -> Mistral SDK is simplest for Mistral-only use

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Using response_format (snake_case) instead of responseFormat (camelCase) -- silently ignored, no error thrown
  • Using input (singular) for embeddings instead of inputs (plural) -- Mistral-specific naming
  • Not configuring retries for production (SDK defaults to strategy: "none" -- zero retries)
  • Hardcoding API keys instead of using environment variables
  • Accessing chunk.choices[0]?.delta?.content directly on streaming events instead of event.data.choices[0]?.delta?.content

Medium Priority Issues:

  • Not setting timeoutMs for production (default is -1, meaning no timeout -- requests can hang indefinitely)
  • Using max_tokens instead of maxTokens (camelCase SDK convention)
  • Missing system role message for behavior guidance
  • Using tool_choice instead of toolChoice
  • Using tool_calls instead of toolCalls when reading responses

Common Mistakes:

  • Importing from "mistralai" instead of "@mistralai/mistralai" -- the correct package name has the org scope
  • Using CommonJS require() -- the package is ESM-only, use import or await import()
  • Confusing Mistral's imageUrl: "url" (flat string) with OpenAI's image_url: { url: "..." } (nested object)
  • Using client.chat.completions.create() (OpenAI pattern) instead of client.chat.complete() (Mistral pattern)
  • Assuming embedding dimensions match OpenAI's -- mistral-embed returns 1024-dimensional vectors, not 1536

Gotchas & Edge Cases:

  • The SDK is ESM-only. In CommonJS projects, you must use const { Mistral } = await import("@mistralai/mistralai").
  • Streaming content may be string | string[] -- cast or check type when writing to stdout.
  • chat.parse() requires a Zod schema passed to responseFormat -- it does not accept { type: "json_object" }.
  • The apiKey constructor option accepts a string OR an async function () => Promise<string> for dynamic key rotation.
  • Model aliases like mistral-large-latest resolve to the latest version of that model tier. Pin to specific versions (e.g., mistral-large-3-25-12) for reproducibility.
  • toolChoice: "any" forces the model to call a tool. toolChoice: "auto" lets the model decide. toolChoice: "none" prevents tool calls.
  • parallelToolCalls: false forces sequential tool calling (default true allows parallel).
  • FIM endpoint (fim.complete()) uses prompt + suffix parameters, NOT the messages array.
  • safePrompt: true injects Mistral's safety system prompt before your messages.
  • The SDK provides standalone functions (e.g., chatComplete() from "@mistralai/mistralai/funcs/chatComplete.js") for tree-shaking in browser/edge runtimes.

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST use responseFormat (camelCase) in SDK calls -- NOT response_format (snake_case). The SDK uses camelCase property names throughout.)

(You MUST configure retries explicitly -- the SDK defaults to strategy: "none" (no retries), unlike OpenAI's SDK which retries automatically)

(You MUST consume streaming results with for await (const event of result) and access content via event.data.choices[0]?.delta?.content -- the event shape differs from OpenAI)

(You MUST never hardcode API keys -- use process.env["MISTRAL_API_KEY"] with the bracket notation the SDK documents)

(You MUST use client.chat.parse() with a Zod schema for structured outputs -- NOT manual JSON.parse() on completion content)

Failure to follow these rules will produce broken API calls (snake_case properties silently ignored), unreliable production services (no retries), or incorrectly parsed streaming data.

</critical_reminders>

Files (skills)
  • examples
    • chat.md 5.7 KB
      # Mistral SDK -- Chat & Streaming Examples
      
      > Chat completions, streaming with async iteration, multi-turn conversations, and token tracking. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [function-calling.md](function-calling.md) -- Tool/function calling
      - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision
      - [codestral.md](codestral.md) -- Codestral FIM code completion
      
      ---
      
      ## Basic Chat Completion
      
      ```typescript
      // basic-chat.ts
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      async function chat(userMessage: string): Promise<string> {
        const result = await client.chat.complete({
          model: "mistral-large-latest",
          messages: [
            { role: "system", content: "You are a helpful assistant. Be concise." },
            { role: "user", content: userMessage },
          ],
        });
      
        const content = result?.choices?.[0]?.message?.content;
        if (!content) {
          throw new Error("No content in response");
        }
      
        // Log token usage
        if (result.usage) {
          console.log(`Tokens: ${result.usage.totalTokens}`);
        }
      
        return typeof content === "string" ? content : content.join("");
      }
      
      const answer = await chat("What is TypeScript in one sentence?");
      console.log(answer);
      ```
      
      **Why good:** Safe optional chaining on nullable fields, handles `content` being `string | string[]`, token tracking
      
      ---
      
      ## Streaming with Async Iteration
      
      ```typescript
      // streaming-chat.ts
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      const result = await client.chat.stream({
        model: "mistral-large-latest",
        messages: [
          { role: "system", content: "You are a helpful assistant." },
          { role: "user", content: "Explain async/await in TypeScript." },
        ],
      });
      
      // IMPORTANT: events wrap data in event.data -- not directly on event
      for await (const event of result) {
        const content = event.data.choices[0]?.delta?.content;
        if (content) {
          // content may be string | string[] depending on model
          process.stdout.write(
            typeof content === "string" ? content : content.join(""),
          );
        }
      }
      console.log(); // newline
      ```
      
      **Why good:** Accesses `event.data.choices[0]` (not `event.choices[0]`), handles string union type
      
      ### BAD: OpenAI-style streaming (wrong for Mistral)
      
      ```typescript
      // BAD: Trying OpenAI's event shape
      for await (const chunk of result) {
        process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); // WRONG: no .data wrapper
      }
      ```
      
      **Why bad:** Mistral wraps streaming data in `event.data` -- accessing `chunk.choices` directly fails silently or throws
      
      ---
      
      ## Multi-Turn Conversation
      
      ```typescript
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      interface ChatMessage {
        role: "system" | "user" | "assistant";
        content: string;
      }
      
      const messages: ChatMessage[] = [
        { role: "system", content: "You are a TypeScript expert." },
        { role: "user", content: "What is a union type?" },
      ];
      
      const first = await client.chat.complete({
        model: "mistral-large-latest",
        messages,
      });
      
      const firstContent = first?.choices?.[0]?.message?.content;
      if (firstContent) {
        // Append assistant response for next turn
        messages.push({
          role: "assistant",
          content:
            typeof firstContent === "string" ? firstContent : firstContent.join(""),
        });
        messages.push({ role: "user", content: "Give me a real-world example." });
      
        const followUp = await client.chat.complete({
          model: "mistral-large-latest",
          messages,
        });
      
        console.log(followUp?.choices?.[0]?.message?.content);
      }
      ```
      
      ---
      
      ## Controlling Output Length and Temperature
      
      ```typescript
      const MAX_TOKENS = 500;
      
      const result = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "Summarize this article." }],
        maxTokens: MAX_TOKENS,
        temperature: 0, // Deterministic output
      });
      
      const finishReason = result?.choices?.[0]?.finishReason;
      if (finishReason === "length") {
        console.warn("Output was truncated -- increase maxTokens");
      }
      ```
      
      **Why good:** Uses `maxTokens` (camelCase), checks `finishReason` (camelCase) for truncation
      
      ---
      
      ## Token Usage Tracking
      
      ```typescript
      const result = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "Hello" }],
      });
      
      const usage = result?.usage;
      if (usage) {
        console.log(`Prompt tokens: ${usage.promptTokens}`);
        console.log(`Completion tokens: ${usage.completionTokens}`);
        console.log(`Total tokens: ${usage.totalTokens}`);
      }
      ```
      
      **Why good:** Uses camelCase properties (`promptTokens`, not `prompt_tokens`)
      
      ---
      
      ## Streaming with Abort
      
      ```typescript
      const controller = new AbortController();
      const ABORT_TIMEOUT_MS = 5_000;
      
      setTimeout(() => controller.abort(), ABORT_TIMEOUT_MS);
      
      try {
        const result = await client.chat.stream(
          {
            model: "mistral-large-latest",
            messages: [{ role: "user", content: "Tell me a long story." }],
          },
          {
            fetchOptions: { signal: controller.signal },
          },
        );
      
        for await (const event of result) {
          const content = event.data.choices[0]?.delta?.content;
          if (content) {
            process.stdout.write(
              typeof content === "string" ? content : content.join(""),
            );
          }
        }
      } catch (error) {
        if (error instanceof Error && error.name === "AbortError") {
          console.log("\nStream aborted");
        } else {
          throw error;
        }
      }
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • codestral.md 4 KB
      # Mistral SDK -- Codestral FIM Examples
      
      > Fill-in-middle code completion and code generation with Codestral. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat completions and streaming
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [function-calling.md](function-calling.md) -- Tool/function calling
      - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision
      
      ---
      
      ## Basic Fill-in-Middle (FIM)
      
      ```typescript
      // fim-completion.ts
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      // FIM: provide code before (prompt) and after (suffix) the cursor
      const result = await client.fim.complete({
        model: "codestral-latest",
        prompt: "function fibonacci(n: number): number {\n  if (n <= 1) return n;\n",
        suffix: "}\n\nconsole.log(fibonacci(10));",
        temperature: 0,
      });
      
      const completion = result.choices?.[0]?.message?.content;
      console.log("Generated code:", completion);
      // The model fills in the gap: "  return fibonacci(n - 1) + fibonacci(n - 2);\n"
      ```
      
      **Why good:** Uses dedicated `fim.complete()` endpoint (not chat), separate `prompt` + `suffix`, deterministic with `temperature: 0`
      
      ---
      
      ## FIM with Stop Sequences
      
      ```typescript
      const result = await client.fim.complete({
        model: "codestral-latest",
        prompt: "class UserService {\n  private users: User[] = [];\n\n  ",
        suffix: "\n\n  getUsers(): User[] {\n    return this.users;\n  }\n}",
        temperature: 0,
        stop: ["\n\n"], // Stop at double newline to prevent over-generation
      });
      
      const code = result.choices?.[0]?.message?.content;
      console.log(code);
      ```
      
      ---
      
      ## FIM vs Chat for Code Generation
      
      Use FIM when you have surrounding code context. Use chat when you need full function generation from a description.
      
      ```typescript
      // FIM: filling a gap in existing code
      const fimResult = await client.fim.complete({
        model: "codestral-latest",
        prompt: "const sorted = array.",
        suffix: ";\nconsole.log(sorted);",
        temperature: 0,
      });
      // Model completes: "sort((a, b) => a - b)"
      
      // Chat: generating from a description
      const chatResult = await client.chat.complete({
        model: "codestral-latest",
        messages: [
          {
            role: "system",
            content:
              "You are a TypeScript code generator. Return only code, no explanations.",
          },
          {
            role: "user",
            content:
              "Write a function that sorts an array of numbers in ascending order.",
          },
        ],
        temperature: 0,
      });
      ```
      
      **When to use FIM:** IDE autocomplete, code insertion at cursor position, completing partial code
      **When to use Chat:** Generating entire functions, explaining code, code review
      
      ---
      
      ## Code Completion with Max Tokens
      
      ```typescript
      const MAX_COMPLETION_TOKENS = 200;
      
      const result = await client.fim.complete({
        model: "codestral-latest",
        prompt: "async function fetchUsers(): Promise<User[]> {\n  ",
        suffix: "\n}",
        temperature: 0,
        maxTokens: MAX_COMPLETION_TOKENS,
      });
      
      const finishReason = result.choices?.[0]?.finishReason;
      if (finishReason === "length") {
        console.warn("Code completion was truncated -- increase maxTokens");
      }
      ```
      
      ---
      
      ## Code Embeddings with Codestral Embed
      
      For code-specific semantic search, use `codestral-embed-latest` instead of `mistral-embed`.
      
      ```typescript
      const CODE_EMBEDDING_MODEL = "codestral-embed-latest";
      
      const result = await client.embeddings.create({
        model: CODE_EMBEDDING_MODEL,
        inputs: [
          "function add(a: number, b: number): number { return a + b; }",
          "const sum = (x: number, y: number): number => x + y;",
          "class Calculator { add(a: number, b: number) { return a + b; } }",
        ],
      });
      
      // Use for code search, duplicate detection, or semantic code navigation
      const vectors = result.data?.map((item) => item.embedding) ?? [];
      console.log(`Generated ${vectors.length} code embeddings`);
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • core.md 6.4 KB
      # Mistral SDK -- Setup & Configuration Examples
      
      > Client initialization, environment config, production settings, error handling, custom HTTP client, and retry configuration. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [chat.md](chat.md) -- Chat completions and streaming
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [function-calling.md](function-calling.md) -- Tool/function calling
      - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision
      - [codestral.md](codestral.md) -- Codestral FIM code completion
      
      ---
      
      ## Basic Client Setup
      
      ```typescript
      // lib/mistral.ts
      import { Mistral } from "@mistralai/mistralai";
      
      // Reads MISTRAL_API_KEY from env
      const client = new Mistral({
        apiKey: process.env["MISTRAL_API_KEY"] ?? "",
      });
      
      export { client };
      ```
      
      ---
      
      ## Production Configuration
      
      ```typescript
      // lib/mistral.ts
      import { Mistral } from "@mistralai/mistralai";
      
      const TIMEOUT_MS = 30_000;
      const INITIAL_RETRY_INTERVAL_MS = 1_000;
      const MAX_RETRY_INTERVAL_MS = 30_000;
      const RETRY_EXPONENT = 1.5;
      const MAX_ELAPSED_TIME_MS = 120_000;
      
      const client = new Mistral({
        apiKey: process.env["MISTRAL_API_KEY"] ?? "",
        timeoutMs: TIMEOUT_MS,
        retryConfig: {
          strategy: "backoff",
          backoff: {
            initialInterval: INITIAL_RETRY_INTERVAL_MS,
            maxInterval: MAX_RETRY_INTERVAL_MS,
            exponent: RETRY_EXPONENT,
            maxElapsedTime: MAX_ELAPSED_TIME_MS,
          },
          retryConnectionErrors: true,
        },
      });
      
      export { client };
      ```
      
      **Why good:** Explicit retry config (SDK defaults to zero retries), named constants, sensible backoff curve
      
      ---
      
      ## Async API Key Provider
      
      ```typescript
      // lib/mistral.ts -- dynamic key rotation
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({
        apiKey: async () => {
          // Fetch from secrets manager, vault, or key rotation service
          return await getSecureApiKey();
        },
        timeoutMs: 30_000,
      });
      
      export { client };
      ```
      
      **Why good:** Supports key rotation without restarting the process -- the SDK calls the function before each request
      
      ---
      
      ## Custom HTTP Client with Hooks
      
      ```typescript
      // lib/mistral.ts -- custom HTTP client
      import { Mistral } from "@mistralai/mistralai";
      import { HTTPClient } from "@mistralai/mistralai/lib/http";
      
      const REQUEST_TIMEOUT_MS = 10_000;
      
      const httpClient = new HTTPClient({
        fetcher: (request) => fetch(request),
      });
      
      // Add timeout to every request
      httpClient.addHook("beforeRequest", (request) => {
        const nextRequest = new Request(request, {
          signal: request.signal || AbortSignal.timeout(REQUEST_TIMEOUT_MS),
        });
        nextRequest.headers.set("x-custom-header", "custom-value");
        return nextRequest;
      });
      
      // Log errors
      httpClient.addHook("requestError", (error, request) => {
        console.error(
          `Mistral request failed: ${request.method} ${request.url}`,
          error,
        );
      });
      
      const client = new Mistral({
        httpClient,
        apiKey: process.env["MISTRAL_API_KEY"] ?? "",
      });
      
      export { client };
      ```
      
      ---
      
      ## Production Error Handling
      
      ```typescript
      // error-handling.ts
      import { Mistral } from "@mistralai/mistralai";
      import {
        SDKError,
        SDKValidationError,
        HTTPValidationError,
      } from "@mistralai/mistralai/models/errors";
      
      const TIMEOUT_MS = 30_000;
      
      const client = new Mistral({
        apiKey: process.env["MISTRAL_API_KEY"] ?? "",
        timeoutMs: TIMEOUT_MS,
        retryConfig: {
          strategy: "backoff",
          backoff: {
            initialInterval: 1_000,
            maxInterval: 30_000,
            exponent: 1.5,
            maxElapsedTime: 60_000,
          },
          retryConnectionErrors: true,
        },
      });
      
      async function safeCompletion(prompt: string): Promise<string | null> {
        try {
          const result = await client.chat.complete({
            model: "mistral-large-latest",
            messages: [
              { role: "system", content: "You are a helpful assistant." },
              { role: "user", content: prompt },
            ],
          });
      
          const content = result?.choices?.[0]?.message?.content;
          if (!content) {
            console.warn("No content in response");
            return null;
          }
      
          // Log token usage for cost tracking
          if (result.usage) {
            console.log(
              `Tokens: prompt=${result.usage.promptTokens}, completion=${result.usage.completionTokens}, total=${result.usage.totalTokens}`,
            );
          }
      
          return typeof content === "string" ? content : content.join("");
        } catch (error) {
          if (error instanceof HTTPValidationError) {
            // 422 -- invalid request shape
            console.error("Validation error:", error.message);
            return null;
          }
      
          if (error instanceof SDKValidationError) {
            // Client-side input validation failure
            console.error("Input validation error:", error.message);
            return null;
          }
      
          if (error instanceof SDKError) {
            // General API error (4xx, 5xx)
            console.error(`API error [${error.statusCode}]: ${error.message}`);
      
            if (error.statusCode === 401) {
              throw new Error(
                "Invalid API key. Check MISTRAL_API_KEY environment variable.",
              );
            }
      
            if (error.statusCode === 429) {
              console.error("Rate limited. Retries exhausted.");
              return null;
            }
      
            return null;
          }
      
          // Unknown errors should be re-thrown
          throw error;
        }
      }
      
      const result = await safeCompletion("Hello!");
      if (result) {
        console.log(result);
      } else {
        console.error("Failed to get completion");
      }
      ```
      
      ---
      
      ## Per-Request Override
      
      ```typescript
      // Override retries and timeout for a single request
      const result = await client.chat.complete(
        {
          model: "mistral-large-latest",
          messages: [{ role: "user", content: "Quick question" }],
        },
        {
          retries: {
            strategy: "backoff",
            backoff: {
              initialInterval: 500,
              maxInterval: 5_000,
              exponent: 2,
              maxElapsedTime: 15_000,
            },
          },
          fetchOptions: {
            signal: AbortSignal.timeout(10_000),
          },
        },
      );
      ```
      
      ---
      
      ## Debug Logging
      
      ```typescript
      // Enable via environment variable
      // MISTRAL_DEBUG=true node app.js
      
      // Or via constructor
      const client = new Mistral({
        apiKey: process.env["MISTRAL_API_KEY"] ?? "",
        debugLogger: console,
      });
      
      // Or with a custom logger
      const client = new Mistral({
        apiKey: process.env["MISTRAL_API_KEY"] ?? "",
        debugLogger: {
          log: (msg: string) => logger.info(msg),
          error: (msg: string) => logger.error(msg),
          warn: (msg: string) => logger.warn(msg),
          debug: (msg: string) => logger.debug(msg),
        },
      });
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • embeddings-vision.md 6 KB
      # Mistral SDK -- Embeddings & Vision Examples
      
      > Embeddings for semantic search and vision/image inputs with vision-capable models. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat completions and streaming
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [function-calling.md](function-calling.md) -- Tool/function calling
      - [codestral.md](codestral.md) -- Codestral FIM code completion
      
      ---
      
      ## Embeddings and Semantic Search
      
      ```typescript
      // embeddings.ts
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      const EMBEDDING_MODEL = "mistral-embed";
      const SIMILARITY_THRESHOLD = 0.7;
      const TOP_K = 3;
      
      function cosineSimilarity(a: number[], b: number[]): number {
        let dot = 0;
        let normA = 0;
        let normB = 0;
        for (let i = 0; i < a.length; i++) {
          dot += a[i] * b[i];
          normA += a[i] * a[i];
          normB += b[i] * b[i];
        }
        return dot / (Math.sqrt(normA) * Math.sqrt(normB));
      }
      
      // Index documents -- batch multiple inputs in one call
      const documents = [
        "TypeScript provides static type checking for JavaScript.",
        "React is a library for building user interfaces.",
        "Node.js is a JavaScript runtime built on V8.",
        "PostgreSQL is a powerful relational database.",
        "Docker containers package applications with dependencies.",
      ];
      
      const docEmbeddings = await client.embeddings.create({
        model: EMBEDDING_MODEL,
        inputs: documents, // NOTE: 'inputs' (plural), not 'input'
      });
      
      const indexedDocs = documents.map((text, i) => ({
        text,
        embedding: docEmbeddings.data?.[i]?.embedding ?? [],
      }));
      
      // Search
      async function search(
        query: string,
      ): Promise<Array<{ text: string; score: number }>> {
        const queryEmbedding = await client.embeddings.create({
          model: EMBEDDING_MODEL,
          inputs: [query],
        });
      
        const queryVector = queryEmbedding.data?.[0]?.embedding ?? [];
      
        return indexedDocs
          .map((doc) => ({
            text: doc.text,
            score: cosineSimilarity(queryVector, doc.embedding),
          }))
          .filter((r) => r.score > SIMILARITY_THRESHOLD)
          .sort((a, b) => b.score - a.score)
          .slice(0, TOP_K);
      }
      
      const results = await search("What is TypeScript?");
      results.forEach((r) => {
        console.log(`[${r.score.toFixed(3)}] ${r.text}`);
      });
      ```
      
      **Why good:** Uses `inputs` (plural, Mistral-specific), batches documents in one call, named constants
      
      ---
      
      ## Embedding Gotchas
      
      ```typescript
      // GOOD: Mistral uses 'inputs' (plural)
      const result = await client.embeddings.create({
        model: "mistral-embed",
        inputs: ["text to embed"],
      });
      // Returns 1024-dimensional vectors
      
      // BAD: OpenAI uses 'input' (singular), dimensions differ
      const result = await client.embeddings.create({
        model: "mistral-embed",
        input: ["text to embed"], // WRONG: use 'inputs'
      });
      // OpenAI returns 1536 dimensions, Mistral returns 1024
      ```
      
      **Why bad:** Wrong parameter name, wrong dimension expectations when migrating from OpenAI
      
      ---
      
      ## Vision -- Image from URL
      
      ```typescript
      // vision.ts
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      async function analyzeImageUrl(
        imageUrl: string,
        question: string,
      ): Promise<string> {
        const result = await client.chat.complete({
          model: "mistral-small-latest", // Vision-capable model
          messages: [
            {
              role: "user",
              content: [
                { type: "text", text: question },
                {
                  type: "image_url",
                  imageUrl: imageUrl, // NOTE: flat string, not { url: "..." }
                },
              ],
            },
          ],
        });
      
        const content = result?.choices?.[0]?.message?.content;
        return typeof content === "string" ? content : (content?.join("") ?? "");
      }
      
      const description = await analyzeImageUrl(
        "https://example.com/photo.jpg",
        "Describe what you see in this image.",
      );
      console.log(description);
      ```
      
      **Why good:** Uses `imageUrl` (camelCase, flat string) -- NOT OpenAI's `image_url: { url: "..." }` nested object
      
      ---
      
      ## Vision -- Local Image (Base64)
      
      ```typescript
      import { Mistral } from "@mistralai/mistralai";
      import { readFileSync } from "node:fs";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      async function analyzeLocalImage(
        imagePath: string,
        question: string,
      ): Promise<string> {
        const imageBuffer = readFileSync(imagePath);
        const base64Image = imageBuffer.toString("base64");
        const mimeType = imagePath.endsWith(".png") ? "image/png" : "image/jpeg";
      
        const result = await client.chat.complete({
          model: "mistral-small-latest",
          messages: [
            {
              role: "user",
              content: [
                { type: "text", text: question },
                {
                  type: "image_url",
                  imageUrl: `data:${mimeType};base64,${base64Image}`,
                },
              ],
            },
          ],
        });
      
        const content = result?.choices?.[0]?.message?.content;
        return typeof content === "string" ? content : (content?.join("") ?? "");
      }
      ```
      
      ---
      
      ## Vision -- Multiple Images
      
      ```typescript
      const result = await client.chat.complete({
        model: "mistral-small-latest",
        messages: [
          {
            role: "user",
            content: [
              { type: "text", text: "Compare these two images." },
              {
                type: "image_url",
                imageUrl: "https://example.com/image1.jpg",
              },
              {
                type: "image_url",
                imageUrl: "https://example.com/image2.jpg",
              },
            ],
          },
        ],
      });
      ```
      
      ---
      
      ## Vision Model Selection
      
      ```
      Vision-capable models (as of 2026):
      - mistral-small-latest (Mistral Small 4)
      - mistral-medium-latest (Mistral Medium 3.1)
      - mistral-large-latest (Mistral Large 3)
      - ministral-14b-latest, ministral-8b-latest, ministral-3b-latest (Ministral 3)
      ```
      
      Most current Mistral models support vision. Use `mistral-small-latest` for cost-efficient image analysis.
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • function-calling.md 6.8 KB
      # Mistral SDK -- Function Calling Examples
      
      > Tool definitions, tool call loop, parallel tool calls, streaming function calling. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat completions and streaming
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision
      - [codestral.md](codestral.md) -- Codestral FIM code completion
      
      ---
      
      ## Basic Function Calling
      
      ```typescript
      // function-calling.ts
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      const tools = [
        {
          type: "function" as const,
          function: {
            name: "get_weather",
            description: "Get the current weather for a location",
            parameters: {
              type: "object",
              properties: {
                location: { type: "string", description: "City name" },
                unit: {
                  type: "string",
                  enum: ["celsius", "fahrenheit"],
                  description: "Temperature unit",
                },
              },
              required: ["location"],
            },
          },
        },
      ];
      
      const result = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "What is the weather in Paris?" }],
        tools,
        toolChoice: "any", // Forces tool use
      });
      
      const toolCall = result?.choices?.[0]?.message?.toolCalls?.[0];
      if (toolCall) {
        const args = JSON.parse(toolCall.function.arguments);
        console.log(`Call ${toolCall.function.name} with:`, args);
        console.log(`Tool call ID: ${toolCall.id}`);
      }
      ```
      
      **Why good:** Uses `toolChoice` (camelCase), `toolCalls` (camelCase), `as const` for type literal, checks for tool calls
      
      ---
      
      ## Complete Tool Call Loop
      
      The standard pattern: send message with tools -> get tool calls -> execute -> send results back.
      
      ```typescript
      // tool-loop.ts
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      // Tool implementations
      async function getWeather(args: {
        location: string;
        unit?: string;
      }): Promise<string> {
        return JSON.stringify({
          location: args.location,
          temperature: 22,
          unit: args.unit ?? "celsius",
          condition: "sunny",
        });
      }
      
      async function searchDatabase(args: { query: string }): Promise<string> {
        return JSON.stringify({
          results: [{ id: 1, title: `Result for: ${args.query}` }],
        });
      }
      
      const toolImplementations: Record<
        string,
        (args: Record<string, unknown>) => Promise<string>
      > = {
        get_weather: getWeather as (args: Record<string, unknown>) => Promise<string>,
        search_database: searchDatabase as (
          args: Record<string, unknown>,
        ) => Promise<string>,
      };
      
      const tools = [
        {
          type: "function" as const,
          function: {
            name: "get_weather",
            description: "Get current weather for a location",
            parameters: {
              type: "object",
              properties: {
                location: { type: "string", description: "City name" },
                unit: { type: "string", enum: ["celsius", "fahrenheit"] },
              },
              required: ["location"],
            },
          },
        },
        {
          type: "function" as const,
          function: {
            name: "search_database",
            description: "Search the knowledge base",
            parameters: {
              type: "object",
              properties: {
                query: { type: "string", description: "Search query" },
              },
              required: ["query"],
            },
          },
        },
      ];
      
      interface ChatMessage {
        role: "system" | "user" | "assistant" | "tool";
        content: string;
        name?: string;
        toolCallId?: string;
        toolCalls?: Array<{
          id: string;
          function: { name: string; arguments: string };
        }>;
      }
      
      const MAX_TOOL_ITERATIONS = 5;
      
      const messages: ChatMessage[] = [
        {
          role: "system",
          content: "You help users with weather and search queries.",
        },
        {
          role: "user",
          content: "What is the weather in London and search for TypeScript guides?",
        },
      ];
      
      let iterations = 0;
      
      while (iterations < MAX_TOOL_ITERATIONS) {
        const result = await client.chat.complete({
          model: "mistral-large-latest",
          messages,
          tools,
        });
      
        const choice = result?.choices?.[0];
        if (!choice) break;
      
        const assistantMessage = choice.message;
      
        // Check if the model wants to call tools
        if (!assistantMessage?.toolCalls || assistantMessage.toolCalls.length === 0) {
          // No tool calls -- model has a final answer
          console.log("Final answer:", assistantMessage?.content);
          break;
        }
      
        // Add assistant message with tool calls to history
        messages.push({
          role: "assistant",
          content:
            typeof assistantMessage.content === "string"
              ? assistantMessage.content
              : "",
          toolCalls: assistantMessage.toolCalls.map((tc) => ({
            id: tc.id ?? "",
            function: { name: tc.function.name, arguments: tc.function.arguments },
          })),
        });
      
        // Execute each tool call and add results
        for (const toolCall of assistantMessage.toolCalls) {
          const fnName = toolCall.function.name;
          const args = JSON.parse(toolCall.function.arguments);
          console.log(`Calling ${fnName}(${JSON.stringify(args)})`);
      
          const impl = toolImplementations[fnName];
          if (!impl) {
            throw new Error(`Unknown tool: ${fnName}`);
          }
      
          const toolResult = await impl(args);
      
          messages.push({
            role: "tool",
            name: fnName,
            content: toolResult,
            toolCallId: toolCall.id ?? "",
          });
        }
      
        iterations++;
      }
      ```
      
      **Why good:** Bounded loop with MAX_TOOL_ITERATIONS, proper tool message format with `toolCallId`, handles parallel tool calls
      
      ---
      
      ## Tool Choice Options
      
      ```typescript
      // "auto" (default) -- model decides whether to call tools
      const result1 = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "Hello" }],
        tools,
        toolChoice: "auto",
      });
      
      // "any" -- forces the model to call at least one tool
      const result2 = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "Get weather in Tokyo" }],
        tools,
        toolChoice: "any",
      });
      
      // "none" -- prevents any tool calls
      const result3 = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "Tell me about weather" }],
        tools,
        toolChoice: "none",
      });
      ```
      
      ---
      
      ## Sequential Tool Calls
      
      ```typescript
      // Force sequential tool calling (disable parallel calls)
      const result = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "Weather in Paris and London?" }],
        tools,
        parallelToolCalls: false, // Forces sequential -- model calls one tool at a time
      });
      ```
      
      **Why good:** `parallelToolCalls: false` ensures tools are called one at a time, useful for dependent operations
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • structured-output.md 5.3 KB
      # Mistral SDK -- Structured Output Examples
      
      > Type-safe structured responses with Zod via `chat.parse()`, and JSON mode via `responseFormat`. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat completions and streaming
      - [function-calling.md](function-calling.md) -- Tool/function calling
      - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision
      - [codestral.md](codestral.md) -- Codestral FIM code completion
      
      ---
      
      ## Structured Output with `chat.parse()` and Zod
      
      ```typescript
      // structured-output.ts
      import { Mistral } from "@mistralai/mistralai";
      import { z } from "zod";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      const ArticleSummary = z.object({
        title: z.string(),
        summary: z.string(),
        keyPoints: z.array(z.string()),
        sentiment: z.enum(["positive", "negative", "neutral"]),
      });
      
      type ArticleSummary = z.infer<typeof ArticleSummary>;
      
      const MAX_TOKENS = 512;
      
      async function extractSummary(
        articleText: string,
      ): Promise<ArticleSummary | null> {
        const result = await client.chat.parse({
          model: "mistral-large-latest",
          messages: [
            {
              role: "system",
              content: "Extract a structured summary from the article.",
            },
            { role: "user", content: articleText },
          ],
          responseFormat: ArticleSummary, // Pass Zod schema directly
          maxTokens: MAX_TOKENS,
          temperature: 0,
        });
      
        // Access the typed parsed result
        const parsed = result.choices?.[0]?.message?.parsed;
        return parsed ?? null;
      }
      
      const article = `
      TypeScript 5.8 brings improved type inference, better error messages,
      and performance optimizations focused on developer experience.
      `;
      
      const summary = await extractSummary(article);
      if (summary) {
        console.log(`Title: ${summary.title}`);
        console.log(`Sentiment: ${summary.sentiment}`);
        console.log("Key Points:");
        summary.keyPoints.forEach((point) => console.log(`  - ${point}`));
      }
      ```
      
      **Why good:** Zod schema passed directly to `responseFormat`, `message.parsed` is fully typed, named constants
      
      ---
      
      ## JSON Mode (Simple)
      
      When you need JSON output without strict schema validation, use `responseFormat: { type: "json_object" }`.
      
      ```typescript
      import { Mistral } from "@mistralai/mistralai";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      const result = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [
          {
            role: "user",
            content:
              "List the top 3 programming languages with their use cases. Return as JSON.",
          },
        ],
        responseFormat: { type: "json_object" },
      });
      
      const content = result?.choices?.[0]?.message?.content;
      if (typeof content === "string") {
        const data = JSON.parse(content);
        console.log(data);
      }
      ```
      
      **Why good:** Uses `responseFormat` (camelCase), explicitly asks for JSON in the prompt (recommended even with JSON mode)
      
      ### BAD: snake_case response_format
      
      ```typescript
      // BAD: Using REST API naming in the SDK
      const result = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "Return JSON" }],
        response_format: { type: "json_object" }, // WRONG: silently ignored
      });
      ```
      
      **Why bad:** The SDK uses camelCase -- `response_format` is silently ignored, model returns plain text
      
      ---
      
      ## Complex Schema with Nested Objects
      
      ```typescript
      import { Mistral } from "@mistralai/mistralai";
      import { z } from "zod";
      
      const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
      
      const Address = z.object({
        street: z.string(),
        city: z.string(),
        country: z.string(),
        postalCode: z.string(),
      });
      
      const Person = z.object({
        name: z.string(),
        age: z.number(),
        email: z.string().email(),
        addresses: z.array(Address),
        occupation: z.string().optional(),
      });
      
      const MAX_TOKENS = 512;
      
      const result = await client.chat.parse({
        model: "mistral-large-latest",
        messages: [
          { role: "system", content: "Extract person details from the text." },
          {
            role: "user",
            content:
              "John Smith, 34, works as a software engineer. " +
              "Email: john@example.com. Lives at 123 Main St, Paris, France 75001.",
          },
        ],
        responseFormat: Person,
        maxTokens: MAX_TOKENS,
        temperature: 0,
      });
      
      const person = result.choices?.[0]?.message?.parsed;
      if (person) {
        console.log(`${person.name}, age ${person.age}`);
        console.log(`Email: ${person.email}`);
        person.addresses.forEach((addr) => {
          console.log(`Address: ${addr.street}, ${addr.city}, ${addr.country}`);
        });
      }
      ```
      
      ---
      
      ## Accessing Raw vs Parsed Content
      
      ```typescript
      const result = await client.chat.parse({
        model: "mistral-large-latest",
        messages: [
          { role: "system", content: "Extract book information." },
          { role: "user", content: "I read 'Dune' by Frank Herbert." },
        ],
        responseFormat: z.object({
          name: z.string(),
          authors: z.array(z.string()),
        }),
      });
      
      const message = result.choices?.[0]?.message;
      
      // Raw JSON string (always available)
      console.log("Raw:", message?.content);
      
      // Parsed and typed object (available via chat.parse)
      console.log("Parsed:", message?.parsed);
      ```
      
      **Why good:** Shows both access paths -- `content` for raw JSON, `parsed` for typed object
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
  • reference.md 9.9 KB
    # Mistral SDK Quick Reference
    
    > Client configuration, model IDs, API methods, error types, and SDK conventions. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Package Installation
    
    ```bash
    # Core package (ESM-only)
    npm install @mistralai/mistralai
    
    # For structured outputs (recommended)
    npm install zod
    ```
    
    ---
    
    ## Client Configuration
    
    ```typescript
    import { Mistral } from "@mistralai/mistralai";
    
    const client = new Mistral({
      apiKey: process.env["MISTRAL_API_KEY"] ?? "", // string or async () => Promise<string>
      timeoutMs: 30_000, // Request timeout in ms (default: -1 = no timeout)
      server: "eu", // Named server selection (default: "eu")
      serverURL: "https://custom-api.example.com", // Custom endpoint override
      retryConfig: {
        // Default: { strategy: "none" } -- NO RETRIES
        strategy: "backoff",
        backoff: {
          initialInterval: 1_000,
          maxInterval: 30_000,
          exponent: 1.5,
          maxElapsedTime: 120_000,
        },
        retryConnectionErrors: true,
      },
      debugLogger: console, // Enable debug logging (or set MISTRAL_DEBUG=true)
    });
    ```
    
    ### Environment Variables
    
    | Variable          | Purpose                       |
    | ----------------- | ----------------------------- |
    | `MISTRAL_API_KEY` | API key (required)            |
    | `MISTRAL_DEBUG`   | Enable debug logging (`true`) |
    
    ### Configuration Priority
    
    ```
    Request-level options (highest)
      -> Client initialization options
        -> Environment variables
          -> SDK defaults (lowest)
    ```
    
    ---
    
    ## Model IDs
    
    ### Language Models (Chat / Text Generation)
    
    | Model ID                   | Alias                   | Use Case                       | Context |
    | -------------------------- | ----------------------- | ------------------------------ | ------- |
    | `mistral-large-3-25-12`    | `mistral-large-latest`  | Most capable, general purpose  | 256K    |
    | `mistral-medium-3-1-25-08` | `mistral-medium-latest` | Balanced cost/performance      | 256K    |
    | `mistral-small-4-0-26-03`  | `mistral-small-latest`  | Cost-efficient, vision-capable | 256K    |
    | `ministral-3-14b-25-12`    | `ministral-14b-latest`  | Compact, multimodal            | 128K    |
    | `ministral-3-8b-25-12`     | `ministral-8b-latest`   | Compact, multimodal            | 128K    |
    | `ministral-3-3b-25-12`     | `ministral-3b-latest`   | Edge / minimal                 | 128K    |
    
    ### Reasoning Models
    
    | Model ID                     | Alias                     | Use Case          |
    | ---------------------------- | ------------------------- | ----------------- |
    | `magistral-medium-1-2-25-09` | `magistral-medium-latest` | Complex reasoning |
    | `magistral-small-1-2-25-09`  | `magistral-small-latest`  | Fast reasoning    |
    
    ### Code Models
    
    | Model ID           | Alias              | Use Case                      |
    | ------------------ | ------------------ | ----------------------------- |
    | `codestral-25-08`  | `codestral-latest` | Code generation + FIM         |
    | `devstral-2-25-12` | `devstral-latest`  | Code agents, codebase explore |
    
    ### Embedding Models
    
    | Model ID                | Alias                    | Dimensions | Max Input   |
    | ----------------------- | ------------------------ | ---------- | ----------- |
    | `mistral-embed-23-12`   | `mistral-embed`          | 1024       | 8192 tokens |
    | `codestral-embed-25-05` | `codestral-embed-latest` | 1024       | 8192 tokens |
    
    ### Specialist Models
    
    | Model ID                        | Use Case            |
    | ------------------------------- | ------------------- |
    | `mistral-ocr-3-25-12`           | Document OCR        |
    | `mistral-moderation-26-03`      | Content moderation  |
    | `voxtral-mini-transcribe-26-02` | Audio transcription |
    
    ---
    
    ## API Methods Reference
    
    ### Chat Completions
    
    ```typescript
    // Standard completion
    const result = await client.chat.complete({
      model: "mistral-large-latest", // Required
      messages: [], // Required: ChatCompletionRequestMessage[]
      temperature: 0.7, // 0.0-1.5 (default: 0.7)
      maxTokens: 1000, // Max output tokens
      topP: 1, // Nucleus sampling (default: 1)
      frequencyPenalty: 0, // Penalize frequent tokens
      presencePenalty: 0, // Penalize present tokens
      tools: [], // Tool definitions
      toolChoice: "auto", // "auto" | "any" | "none" | "required"
      parallelToolCalls: true, // Allow parallel tool calls
      responseFormat: { type: "text" }, // "text" | "json_object"
      safePrompt: false, // Inject safety prompt
      stop: undefined, // string | string[] -- stop sequences
    });
    
    // Structured output parsing with Zod
    const result = await client.chat.parse({
      model: "mistral-large-latest",
      messages: [],
      responseFormat: zodSchema, // Pass Zod schema directly
      maxTokens: 256,
      temperature: 0,
    });
    
    // Streaming
    const result = await client.chat.stream({
      model: "mistral-large-latest",
      messages: [],
    });
    
    for await (const event of result) {
      // event.data.choices[0]?.delta?.content
    }
    ```
    
    ### FIM (Fill-in-Middle)
    
    ```typescript
    const result = await client.fim.complete({
      model: "codestral-latest", // Required
      prompt: "", // Code before cursor (required)
      suffix: "", // Code after cursor (optional)
      temperature: 0, // Sampling temperature
      maxTokens: 1000, // Max output tokens
      stop: undefined, // Stop sequences
    });
    ```
    
    ### Embeddings
    
    ```typescript
    const result = await client.embeddings.create({
      model: "mistral-embed", // Required
      inputs: [], // string[] (NOTE: plural 'inputs', not 'input')
    });
    ```
    
    ### Agents
    
    ```typescript
    const result = await client.agents.complete({
      agentId: "<id>", // Required: pre-configured agent ID
      messages: [], // Required
      responseFormat: { type: "text" },
    });
    
    // Streaming
    const result = await client.agents.stream({
      agentId: "<id>",
      messages: [],
    });
    ```
    
    ### Files
    
    ```typescript
    import { openAsBlob } from "node:fs";
    
    // Upload
    const result = await client.files.upload({
      file: await openAsBlob("data.jsonl"),
    });
    
    // List, retrieve, delete
    const files = await client.files.list();
    const file = await client.files.retrieve({ fileId: "file-abc123" });
    await client.files.delete({ fileId: "file-abc123" });
    ```
    
    ### Models
    
    ```typescript
    const models = await client.models.list();
    const model = await client.models.retrieve({ modelId: "mistral-large-latest" });
    ```
    
    ---
    
    ## Error Types
    
    | Error Class           | Description                                |
    | --------------------- | ------------------------------------------ |
    | `SDKError`            | General API errors (4XX, 5XX status codes) |
    | `SDKValidationError`  | Client-side input validation failure       |
    | `HTTPValidationError` | Server-side validation (422)               |
    | `ConnectionError`     | Network connectivity issues                |
    | `RequestTimeoutError` | Request timeout exceeded                   |
    | `RequestAbortedError` | Client-cancelled request                   |
    
    ```typescript
    import {
      SDKError,
      SDKValidationError,
      HTTPValidationError,
    } from "@mistralai/mistralai/models/errors";
    ```
    
    ---
    
    ## Retry Configuration
    
    ```typescript
    // Global retry (at client init)
    retryConfig: {
      strategy: "backoff", // "none" | "backoff"
      backoff: {
        initialInterval: 1_000,  // ms between first retry
        maxInterval: 30_000,     // ms max between retries
        exponent: 1.5,           // backoff multiplier
        maxElapsedTime: 120_000, // ms total retry window
      },
      retryConnectionErrors: true,
    }
    
    // Per-request retry override
    await client.chat.complete(
      { model: "mistral-large-latest", messages: [...] },
      {
        retries: {
          strategy: "backoff",
          backoff: { initialInterval: 500, maxInterval: 5_000, exponent: 2, maxElapsedTime: 30_000 },
        },
      },
    );
    ```
    
    ---
    
    ## Custom HTTP Client
    
    ```typescript
    import { HTTPClient } from "@mistralai/mistralai/lib/http";
    
    const httpClient = new HTTPClient({
      fetcher: (request) => fetch(request),
    });
    
    httpClient.addHook("beforeRequest", (request) => {
      const nextRequest = new Request(request, {
        signal: request.signal || AbortSignal.timeout(5_000),
      });
      nextRequest.headers.set("x-custom-header", "value");
      return nextRequest;
    });
    
    httpClient.addHook("requestError", (error, request) => {
      console.error(`Request failed: ${request.method} ${request.url}`, error);
    });
    
    const client = new Mistral({
      httpClient,
      apiKey: process.env["MISTRAL_API_KEY"] ?? "",
    });
    ```
    
    ---
    
    ## SDK vs REST API Property Names
    
    The SDK uses camelCase. The REST API uses snake_case. This table maps the most common properties:
    
    | SDK (camelCase)     | REST API (snake_case) |
    | ------------------- | --------------------- |
    | `responseFormat`    | `response_format`     |
    | `maxTokens`         | `max_tokens`          |
    | `topP`              | `top_p`               |
    | `toolChoice`        | `tool_choice`         |
    | `toolCalls`         | `tool_calls`          |
    | `parallelToolCalls` | `parallel_tool_calls` |
    | `frequencyPenalty`  | `frequency_penalty`   |
    | `presencePenalty`   | `presence_penalty`    |
    | `safePrompt`        | `safe_prompt`         |
    | `imageUrl`          | `image_url`           |
    | `agentId`           | `agent_id`            |
    
    ---
    
    ## Standalone Functions (Tree-Shaking)
    
    For browser/edge runtimes, import standalone functions to reduce bundle size:
    
    ```typescript
    import { chatComplete } from "@mistralai/mistralai/funcs/chatComplete.js";
    import { chatStream } from "@mistralai/mistralai/funcs/chatStream.js";
    import { embeddingsCreate } from "@mistralai/mistralai/funcs/embeddingsCreate.js";
    
    const res = await chatComplete(client, {
      model: "mistral-small-latest",
      messages: [{ role: "user", content: "Hello" }],
    });
    
    if (res.ok) {
      console.log(res.value);
    }
    ```
    
    ---
    
    ## Message Roles
    
    | Role        | Description                                            |
    | ----------- | ------------------------------------------------------ |
    | `system`    | System instruction (behavior guidance)                 |
    | `user`      | User input                                             |
    | `assistant` | Model response                                         |
    | `tool`      | Tool result (requires `name`, `content`, `toolCallId`) |
    
  • SKILL.md 21.5 KB
    ---
    name: ai-provider-mistral-sdk
    description: Official Mistral AI TypeScript SDK patterns — client setup, chat completions, streaming, function calling, structured outputs, embeddings, vision, Codestral FIM, and production best practices
    ---
    
    # Mistral SDK Patterns
    
    > **Quick Guide:** Use `@mistralai/mistralai` (ESM-only) to interact with Mistral's API. Use `client.chat.complete()` for chat, `client.chat.stream()` for streaming (async iterable via `for await`), `client.chat.parse()` with a Zod schema for structured outputs, and `client.fim.complete()` for Codestral fill-in-middle code completion. The SDK uses `responseFormat` (camelCase) not `response_format`. Streaming events expose content via `event.data.choices[0]?.delta?.content`. Retries default to `strategy: "none"` -- you must configure them explicitly for production.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST use `responseFormat` (camelCase) in SDK calls -- NOT `response_format` (snake_case). The SDK uses camelCase property names throughout.)**
    
    **(You MUST configure retries explicitly -- the SDK defaults to `strategy: "none"` (no retries), unlike OpenAI's SDK which retries automatically)**
    
    **(You MUST consume streaming results with `for await (const event of result)` and access content via `event.data.choices[0]?.delta?.content` -- the event shape differs from OpenAI)**
    
    **(You MUST never hardcode API keys -- use `process.env["MISTRAL_API_KEY"]` with the bracket notation the SDK documents)**
    
    **(You MUST use `client.chat.parse()` with a Zod schema for structured outputs -- NOT manual `JSON.parse()` on completion content)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** Mistral, mistral, @mistralai/mistralai, client.chat.complete, client.chat.stream, client.chat.parse, client.fim.complete, client.embeddings.create, mistral-large, mistral-small, codestral, pixtral, ministral, magistral, devstral, MISTRAL_API_KEY, responseFormat, mistral-embed
    
    **When to use:**
    
    - Building applications that call Mistral models directly (Mistral Large, Small, Codestral, etc.)
    - Implementing chat completions with SSE streaming
    - Using Codestral for code generation and fill-in-middle (FIM) completion
    - Extracting structured data with `client.chat.parse()` and Zod schemas
    - Implementing function calling / tool use
    - Creating embeddings for RAG pipelines or semantic search
    - Processing images with vision-capable models (Mistral Small, Medium, Large, Ministral)
    - Using Mistral Agents API for pre-configured agent completions
    
    **Key patterns covered:**
    
    - Client initialization and configuration (retries, timeouts, custom HTTP client)
    - Chat completions (`chat.complete`) and streaming (`chat.stream`)
    - Structured outputs with `chat.parse()` and Zod schemas
    - Function calling / tool use with tool call loop
    - Embeddings (`embeddings.create`) with `mistral-embed`
    - Vision (image URL / base64 with vision-capable models)
    - Codestral FIM (`fim.complete`) for code completion
    - Error handling, retry configuration, and production patterns
    
    **When NOT to use:**
    
    - Multi-provider applications where you need to switch between Mistral, OpenAI, Anthropic, etc. -- use a unified provider SDK
    - React-specific chat UI hooks (`useChat`) -- use a framework-integrated AI SDK
    - When you need OpenAI-compatible endpoints -- use OpenAI SDK with Mistral's compatible endpoint instead
    
    ---
    
    ## Examples Index
    
    - [Core: Setup & Configuration](examples/core.md) -- Client init, production config, error handling, retries, custom HTTP client
    - [Chat & Streaming](examples/chat.md) -- Chat completions, streaming with async iteration, multi-turn
    - [Structured Output](examples/structured-output.md) -- `chat.parse()` with Zod, JSON mode, typed responses
    - [Function Calling](examples/function-calling.md) -- Tool definitions, tool call loop, streaming tools
    - [Embeddings & Vision](examples/embeddings-vision.md) -- Semantic search, image analysis with vision-capable models
    - [Codestral FIM](examples/codestral.md) -- Fill-in-middle code completion, code generation
    - [Quick API Reference](reference.md) -- Model IDs, method signatures, error types, configuration options
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    The `@mistralai/mistralai` SDK is **auto-generated from Mistral's OpenAPI spec using Speakeasy**, giving you a thin, type-safe wrapper over the REST API. It is ESM-only and uses camelCase property names (not snake_case like the REST API).
    
    **Core principles:**
    
    1. **ESM-only** -- The package is published as ESM only. CommonJS projects must use `await import()`. This is a hard constraint, not optional.
    2. **camelCase API surface** -- SDK properties use camelCase (`responseFormat`, `maxTokens`, `toolChoice`) even though the REST API uses snake_case. This catches OpenAI SDK migrants who write `response_format`.
    3. **No automatic retries** -- Unlike OpenAI's SDK (2 retries by default), Mistral defaults to `strategy: "none"`. You must configure retries explicitly for production.
    4. **Streaming via async iterables** -- `chat.stream()` returns an `EventStream` consumed with `for await...of`. Events have a `data` wrapper: `event.data.choices[0]?.delta?.content`.
    5. **Structured outputs via `chat.parse()`** -- Pass a Zod schema directly to `responseFormat` and access `message.parsed` for typed results. No manual JSON schema construction needed.
    6. **Codestral FIM** -- Dedicated `fim.complete()` endpoint for fill-in-middle code completion, separate from chat.
    
    **When to use the Mistral SDK directly:**
    
    - You only use Mistral models and want the simplest, most direct integration
    - You need Mistral-specific features (Codestral FIM, Mistral Agents, Voxtral audio)
    - You want minimal dependencies and zero abstraction overhead
    - You need the latest Mistral API features on day one
    
    **When NOT to use:**
    
    - You need to switch between providers (OpenAI, Anthropic, Mistral) -- use a unified provider SDK
    - You want React-specific chat UI hooks -- use a framework-integrated AI SDK
    - You want an OpenAI-compatible wrapper -- Mistral exposes an OpenAI-compatible endpoint, use the OpenAI SDK for that
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Client Setup
    
    Initialize the Mistral client. It reads `MISTRAL_API_KEY` from the environment.
    
    ```typescript
    // lib/mistral.ts -- basic setup
    import { Mistral } from "@mistralai/mistralai";
    
    const client = new Mistral({
      apiKey: process.env["MISTRAL_API_KEY"] ?? "",
    });
    
    export { client };
    ```
    
    ```typescript
    // lib/mistral.ts -- production configuration
    import { Mistral } from "@mistralai/mistralai";
    
    const TIMEOUT_MS = 30_000;
    
    const client = new Mistral({
      apiKey: process.env["MISTRAL_API_KEY"] ?? "",
      timeoutMs: TIMEOUT_MS,
      retryConfig: {
        strategy: "backoff",
        backoff: {
          initialInterval: 1_000,
          maxInterval: 30_000,
          exponent: 1.5,
          maxElapsedTime: 120_000,
        },
        retryConnectionErrors: true,
      },
    });
    
    export { client };
    ```
    
    **Why good:** Explicit retry config (SDK defaults to no retries), named constants, env var with bracket notation
    
    **See:** [examples/core.md](examples/core.md) for custom HTTP client, async API key provider, error handling
    
    ---
    
    ### Pattern 2: Chat Completions
    
    Basic chat using `chat.complete()`.
    
    ```typescript
    const result = await client.chat.complete({
      model: "mistral-large-latest",
      messages: [
        { role: "system", content: "You are a helpful coding assistant." },
        { role: "user", content: "Explain TypeScript generics." },
      ],
    });
    
    const content = result?.choices?.[0]?.message?.content;
    console.log(content);
    ```
    
    **Why good:** Uses `system` role for instructions, safe optional chaining on nullable response
    
    ```typescript
    // BAD: Using snake_case properties (REST API style, not SDK style)
    const result = await client.chat.complete({
      model: "mistral-large-latest",
      messages: [{ role: "user", content: "hello" }],
      response_format: { type: "json_object" }, // WRONG: use responseFormat
      max_tokens: 100, // WRONG: use maxTokens
    });
    ```
    
    **Why bad:** SDK uses camelCase properties -- `response_format` and `max_tokens` will be silently ignored
    
    **See:** [examples/chat.md](examples/chat.md) for multi-turn, token tracking, temperature control
    
    ---
    
    ### Pattern 3: Streaming
    
    Use `chat.stream()` for streaming. Events are async iterables.
    
    ```typescript
    const result = await client.chat.stream({
      model: "mistral-large-latest",
      messages: [
        { role: "system", content: "You are a helpful assistant." },
        { role: "user", content: "Explain async/await in TypeScript." },
      ],
    });
    
    for await (const event of result) {
      const content = event.data.choices[0]?.delta?.content;
      if (content) {
        process.stdout.write(content as string);
      }
    }
    console.log();
    ```
    
    **Why good:** Proper `for await` iteration, accesses `event.data` (not `event` directly), handles nullable delta
    
    ```typescript
    // BAD: Trying to access content directly on event (OpenAI pattern)
    for await (const chunk of result) {
      process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); // WRONG
    }
    ```
    
    **Why bad:** Mistral streaming events wrap data in `event.data` -- direct access on `chunk` will fail
    
    **See:** [examples/chat.md](examples/chat.md) for complete streaming examples
    
    ---
    
    ### Pattern 4: Structured Outputs with Zod
    
    Use `chat.parse()` with a Zod schema for type-safe structured responses.
    
    ```typescript
    import { Mistral } from "@mistralai/mistralai";
    import { z } from "zod";
    
    const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" });
    
    const BookSchema = z.object({
      name: z.string(),
      authors: z.array(z.string()),
    });
    
    const MAX_TOKENS = 256;
    
    const result = await client.chat.parse({
      model: "mistral-large-latest",
      messages: [
        { role: "system", content: "Extract the book information." },
        { role: "user", content: "I recently read 'Dune' by Frank Herbert." },
      ],
      responseFormat: BookSchema,
      maxTokens: MAX_TOKENS,
      temperature: 0,
    });
    
    const parsed = result.choices?.[0]?.message?.parsed;
    // parsed is typed as { name: string; authors: string[] }
    ```
    
    **Why good:** Schema passed directly to `responseFormat`, `message.parsed` is fully typed, named constants
    
    **See:** [examples/structured-output.md](examples/structured-output.md) for JSON mode, complex schemas
    
    ---
    
    ### Pattern 5: Function Calling / Tool Use
    
    Define tools and handle the tool call loop.
    
    ```typescript
    const tools = [
      {
        type: "function" as const,
        function: {
          name: "get_weather",
          description: "Get current weather for a city",
          parameters: {
            type: "object",
            properties: {
              location: { type: "string", description: "City name" },
            },
            required: ["location"],
          },
        },
      },
    ];
    
    const result = await client.chat.complete({
      model: "mistral-large-latest",
      messages: [{ role: "user", content: "Weather in Paris?" }],
      tools,
      toolChoice: "any",
    });
    
    const toolCall = result?.choices?.[0]?.message?.toolCalls?.[0];
    if (toolCall) {
      const args = JSON.parse(toolCall.function.arguments);
      console.log(`Call ${toolCall.function.name} with:`, args);
    }
    ```
    
    **Why good:** Uses `toolChoice` (camelCase), `toolCalls` (camelCase), proper `as const` for type literal
    
    **See:** [examples/function-calling.md](examples/function-calling.md) for complete tool loop, parallel calls
    
    ---
    
    ### Pattern 6: Embeddings
    
    Create embeddings with `mistral-embed`. Note: uses `inputs` (plural), not `input`.
    
    ```typescript
    const EMBEDDING_MODEL = "mistral-embed";
    
    const result = await client.embeddings.create({
      model: EMBEDDING_MODEL,
      inputs: ["First document", "Second document", "Third document"],
    });
    
    const vectors = result.data?.map((item) => item.embedding) ?? [];
    ```
    
    **Why good:** Uses `inputs` (Mistral-specific, plural), named model constant, safe optional chaining
    
    ```typescript
    // BAD: Using singular 'input' (OpenAI pattern)
    const result = await client.embeddings.create({
      model: "mistral-embed",
      input: ["First document"], // WRONG: Mistral uses 'inputs' (plural)
    });
    ```
    
    **Why bad:** Mistral SDK uses `inputs` (plural) -- `input` (singular) will error or be silently ignored
    
    **See:** [examples/embeddings-vision.md](examples/embeddings-vision.md) for cosine similarity, semantic search
    
    ---
    
    ### Pattern 7: Vision
    
    Send images to vision-capable models using multi-part content arrays.
    
    ```typescript
    const result = await client.chat.complete({
      model: "mistral-small-latest",
      messages: [
        {
          role: "user",
          content: [
            { type: "text", text: "What is in this image?" },
            {
              type: "image_url",
              imageUrl: "https://example.com/photo.jpg",
            },
          ],
        },
      ],
    });
    ```
    
    **Why good:** Uses `imageUrl` (camelCase string), not `image_url: { url }` (OpenAI's nested object pattern)
    
    **See:** [examples/embeddings-vision.md](examples/embeddings-vision.md) for base64 images, multiple images
    
    ---
    
    ### Pattern 8: Codestral FIM
    
    Fill-in-middle code completion using the dedicated FIM endpoint.
    
    ```typescript
    const result = await client.fim.complete({
      model: "codestral-latest",
      prompt: "function fibonacci(n: number): number {\n  if (n <= 1) return n;\n",
      suffix: "}\n\nconsole.log(fibonacci(10));",
      temperature: 0,
    });
    
    const completion = result.choices?.[0]?.message?.content;
    // completion fills the gap between prompt and suffix
    ```
    
    **Why good:** Dedicated FIM endpoint, separate `prompt` + `suffix` (not messages), deterministic with `temperature: 0`
    
    **See:** [examples/codestral.md](examples/codestral.md) for code generation patterns
    
    ---
    
    ### Pattern 9: Error Handling
    
    The SDK provides specific error types. Configure retries since the default is no retries.
    
    ```typescript
    import { Mistral } from "@mistralai/mistralai";
    import {
      SDKError,
      SDKValidationError,
      HTTPValidationError,
    } from "@mistralai/mistralai/models/errors";
    
    try {
      const result = await client.chat.complete({
        model: "mistral-large-latest",
        messages: [{ role: "user", content: "Hello" }],
      });
    } catch (error) {
      if (error instanceof HTTPValidationError) {
        console.error("Validation error:", error.message);
      } else if (error instanceof SDKValidationError) {
        console.error("Input validation error:", error.message);
      } else if (error instanceof SDKError) {
        console.error(`API error [${error.statusCode}]: ${error.message}`);
      } else {
        throw error;
      }
    }
    ```
    
    **Why good:** Specific error types checked in order of specificity, re-throws unexpected errors
    
    **See:** [examples/core.md](examples/core.md) for full error handling, timeout handling, retry configuration
    
    </patterns>
    
    ---
    
    <performance>
    
    ## Performance Optimization
    
    ### Model Selection
    
    ```
    General purpose (most capable)    -> mistral-large-latest (Mistral Large 3)
    Balanced cost/quality             -> mistral-medium-latest (Mistral Medium 3.1)
    Cost-sensitive / fast             -> mistral-small-latest (Mistral Small 4)
    Edge / minimal                    -> ministral-3b-latest or ministral-8b-latest
    Complex reasoning                 -> magistral-medium-latest
    Code generation                   -> codestral-latest or devstral-latest
    Code completion (FIM)             -> codestral-latest (dedicated FIM endpoint)
    Vision / images                   -> mistral-small-latest or mistral-large-latest
    Embeddings                        -> mistral-embed (1024 dimensions)
    Code embeddings                   -> codestral-embed-latest
    ```
    
    ### Key Optimization Patterns
    
    - **Configure retries** -- Default is no retries. Always set `retryConfig` for production.
    - **Set timeouts** -- Default is no timeout (`-1`). Set `timeoutMs` to avoid hanging requests.
    - **Use `temperature: 0`** for deterministic output (enables server-side caching).
    - **Batch embedding inputs** -- Pass multiple strings to `inputs` array in one call.
    - **Use FIM for code completion** -- `fim.complete()` is purpose-built and more efficient than chat for code completion tasks.
    
    </performance>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Method to Use
    
    ```
    What do you need?
    +-- Chat completion (text in, text out)?
    |   +-- Need streaming? -> client.chat.stream()
    |   +-- Need structured JSON? -> client.chat.parse() with Zod schema
    |   +-- Basic completion? -> client.chat.complete()
    +-- Code completion / fill-in-middle?
    |   +-- YES -> client.fim.complete() with Codestral
    +-- Embeddings for search/RAG?
    |   +-- YES -> client.embeddings.create() with mistral-embed
    +-- Pre-configured agent?
        +-- YES -> client.agents.complete() with agent ID
    ```
    
    ### Which Model to Choose
    
    ```
    What is your task?
    +-- Most capable general purpose -> mistral-large-latest
    +-- Balanced cost/performance -> mistral-medium-latest
    +-- Fast + cost-efficient -> mistral-small-latest
    +-- Minimal / edge deployment -> ministral-3b-latest
    +-- Complex reasoning / math -> magistral-medium-latest
    +-- Code generation (chat) -> codestral-latest or devstral-latest
    +-- Code completion (FIM) -> codestral-latest
    +-- Vision / image analysis -> mistral-small-latest (or any vision-capable model)
    +-- Embeddings -> mistral-embed
    +-- Code embeddings -> codestral-embed-latest
    ```
    
    ### Streaming vs Non-Streaming
    
    ```
    Is the response user-facing?
    +-- YES -> Use client.chat.stream()
    |   +-- Iterate with: for await (const event of result)
    |   +-- Access content: event.data.choices[0]?.delta?.content
    +-- NO -> Use client.chat.complete()
        +-- Background processing -> chat.complete()
        +-- Structured output -> chat.parse() with Zod
    ```
    
    ### When to Use This SDK vs a Provider-Agnostic SDK
    
    ```
    Do you need multiple LLM providers (Mistral + others)?
    +-- YES -> Not this skill's scope -- use a unified provider SDK
    +-- NO -> Do you need Mistral-specific features?
        +-- YES -> Use Mistral SDK directly
        |   Examples: Codestral FIM, Mistral Agents,
        |   Voxtral audio, OCR, custom endpoints
        +-- NO -> Mistral SDK is simplest for Mistral-only use
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Using `response_format` (snake_case) instead of `responseFormat` (camelCase) -- silently ignored, no error thrown
    - Using `input` (singular) for embeddings instead of `inputs` (plural) -- Mistral-specific naming
    - Not configuring retries for production (SDK defaults to `strategy: "none"` -- zero retries)
    - Hardcoding API keys instead of using environment variables
    - Accessing `chunk.choices[0]?.delta?.content` directly on streaming events instead of `event.data.choices[0]?.delta?.content`
    
    **Medium Priority Issues:**
    
    - Not setting `timeoutMs` for production (default is `-1`, meaning no timeout -- requests can hang indefinitely)
    - Using `max_tokens` instead of `maxTokens` (camelCase SDK convention)
    - Missing `system` role message for behavior guidance
    - Using `tool_choice` instead of `toolChoice`
    - Using `tool_calls` instead of `toolCalls` when reading responses
    
    **Common Mistakes:**
    
    - Importing from `"mistralai"` instead of `"@mistralai/mistralai"` -- the correct package name has the org scope
    - Using CommonJS `require()` -- the package is ESM-only, use `import` or `await import()`
    - Confusing Mistral's `imageUrl: "url"` (flat string) with OpenAI's `image_url: { url: "..." }` (nested object)
    - Using `client.chat.completions.create()` (OpenAI pattern) instead of `client.chat.complete()` (Mistral pattern)
    - Assuming embedding dimensions match OpenAI's -- `mistral-embed` returns 1024-dimensional vectors, not 1536
    
    **Gotchas & Edge Cases:**
    
    - The SDK is ESM-only. In CommonJS projects, you must use `const { Mistral } = await import("@mistralai/mistralai")`.
    - Streaming content may be `string | string[]` -- cast or check type when writing to stdout.
    - `chat.parse()` requires a Zod schema passed to `responseFormat` -- it does not accept `{ type: "json_object" }`.
    - The `apiKey` constructor option accepts a string OR an async function `() => Promise<string>` for dynamic key rotation.
    - Model aliases like `mistral-large-latest` resolve to the latest version of that model tier. Pin to specific versions (e.g., `mistral-large-3-25-12`) for reproducibility.
    - `toolChoice: "any"` forces the model to call a tool. `toolChoice: "auto"` lets the model decide. `toolChoice: "none"` prevents tool calls.
    - `parallelToolCalls: false` forces sequential tool calling (default `true` allows parallel).
    - FIM endpoint (`fim.complete()`) uses `prompt` + `suffix` parameters, NOT the `messages` array.
    - `safePrompt: true` injects Mistral's safety system prompt before your messages.
    - The SDK provides standalone functions (e.g., `chatComplete()` from `"@mistralai/mistralai/funcs/chatComplete.js"`) for tree-shaking in browser/edge runtimes.
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST use `responseFormat` (camelCase) in SDK calls -- NOT `response_format` (snake_case). The SDK uses camelCase property names throughout.)**
    
    **(You MUST configure retries explicitly -- the SDK defaults to `strategy: "none"` (no retries), unlike OpenAI's SDK which retries automatically)**
    
    **(You MUST consume streaming results with `for await (const event of result)` and access content via `event.data.choices[0]?.delta?.content` -- the event shape differs from OpenAI)**
    
    **(You MUST never hardcode API keys -- use `process.env["MISTRAL_API_KEY"]` with the bracket notation the SDK documents)**
    
    **(You MUST use `client.chat.parse()` with a Zod schema for structured outputs -- NOT manual `JSON.parse()` on completion content)**
    
    **Failure to follow these rules will produce broken API calls (snake_case properties silently ignored), unreliable production services (no retries), or incorrectly parsed streaming data.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related