ai-provider-mistral-sdk
Official Mistral AI TypeScript SDK patterns — client setup, chat completions, streaming, function calling, structured outputs, embeddings, vision, Codestral FIM, and production best practices
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-provider-mistral-sdk/skills/ai-provider-mistral-sdk
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Mistral SDK Patterns
Quick Guide: Use
@mistralai/mistralai(ESM-only) to interact with Mistral's API. Useclient.chat.complete()for chat,client.chat.stream()for streaming (async iterable viafor await),client.chat.parse()with a Zod schema for structured outputs, andclient.fim.complete()for Codestral fill-in-middle code completion. The SDK usesresponseFormat(camelCase) notresponse_format. Streaming events expose content viaevent.data.choices[0]?.delta?.content. Retries default tostrategy: "none"-- you must configure them explicitly for production.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST use responseFormat (camelCase) in SDK calls -- NOT response_format (snake_case). The SDK uses camelCase property names throughout.)
(You MUST configure retries explicitly -- the SDK defaults to strategy: "none" (no retries), unlike OpenAI's SDK which retries automatically)
(You MUST consume streaming results with for await (const event of result) and access content via event.data.choices[0]?.delta?.content -- the event shape differs from OpenAI)
(You MUST never hardcode API keys -- use process.env["MISTRAL_API_KEY"] with the bracket notation the SDK documents)
(You MUST use client.chat.parse() with a Zod schema for structured outputs -- NOT manual JSON.parse() on completion content)
</critical_requirements>
Auto-detection: Mistral, mistral, @mistralai/mistralai, client.chat.complete, client.chat.stream, client.chat.parse, client.fim.complete, client.embeddings.create, mistral-large, mistral-small, codestral, pixtral, ministral, magistral, devstral, MISTRAL_API_KEY, responseFormat, mistral-embed
When to use:
- Building applications that call Mistral models directly (Mistral Large, Small, Codestral, etc.)
- Implementing chat completions with SSE streaming
- Using Codestral for code generation and fill-in-middle (FIM) completion
- Extracting structured data with
client.chat.parse()and Zod schemas - Implementing function calling / tool use
- Creating embeddings for RAG pipelines or semantic search
- Processing images with vision-capable models (Mistral Small, Medium, Large, Ministral)
- Using Mistral Agents API for pre-configured agent completions
Key patterns covered:
- Client initialization and configuration (retries, timeouts, custom HTTP client)
- Chat completions (
chat.complete) and streaming (chat.stream) - Structured outputs with
chat.parse()and Zod schemas - Function calling / tool use with tool call loop
- Embeddings (
embeddings.create) withmistral-embed - Vision (image URL / base64 with vision-capable models)
- Codestral FIM (
fim.complete) for code completion - Error handling, retry configuration, and production patterns
When NOT to use:
- Multi-provider applications where you need to switch between Mistral, OpenAI, Anthropic, etc. -- use a unified provider SDK
- React-specific chat UI hooks (
useChat) -- use a framework-integrated AI SDK - When you need OpenAI-compatible endpoints -- use OpenAI SDK with Mistral's compatible endpoint instead
Examples Index
- Core: Setup & Configuration -- Client init, production config, error handling, retries, custom HTTP client
- Chat & Streaming -- Chat completions, streaming with async iteration, multi-turn
- Structured Output --
chat.parse()with Zod, JSON mode, typed responses - Function Calling -- Tool definitions, tool call loop, streaming tools
- Embeddings & Vision -- Semantic search, image analysis with vision-capable models
- Codestral FIM -- Fill-in-middle code completion, code generation
- Quick API Reference -- Model IDs, method signatures, error types, configuration options
<decision_framework>
Decision Framework
Which Method to Use
What do you need?
+-- Chat completion (text in, text out)?
| +-- Need streaming? -> client.chat.stream()
| +-- Need structured JSON? -> client.chat.parse() with Zod schema
| +-- Basic completion? -> client.chat.complete()
+-- Code completion / fill-in-middle?
| +-- YES -> client.fim.complete() with Codestral
+-- Embeddings for search/RAG?
| +-- YES -> client.embeddings.create() with mistral-embed
+-- Pre-configured agent?
+-- YES -> client.agents.complete() with agent ID
Which Model to Choose
What is your task?
+-- Most capable general purpose -> mistral-large-latest
+-- Balanced cost/performance -> mistral-medium-latest
+-- Fast + cost-efficient -> mistral-small-latest
+-- Minimal / edge deployment -> ministral-3b-latest
+-- Complex reasoning / math -> magistral-medium-latest
+-- Code generation (chat) -> codestral-latest or devstral-latest
+-- Code completion (FIM) -> codestral-latest
+-- Vision / image analysis -> mistral-small-latest (or any vision-capable model)
+-- Embeddings -> mistral-embed
+-- Code embeddings -> codestral-embed-latest
Streaming vs Non-Streaming
Is the response user-facing?
+-- YES -> Use client.chat.stream()
| +-- Iterate with: for await (const event of result)
| +-- Access content: event.data.choices[0]?.delta?.content
+-- NO -> Use client.chat.complete()
+-- Background processing -> chat.complete()
+-- Structured output -> chat.parse() with Zod
When to Use This SDK vs a Provider-Agnostic SDK
Do you need multiple LLM providers (Mistral + others)?
+-- YES -> Not this skill's scope -- use a unified provider SDK
+-- NO -> Do you need Mistral-specific features?
+-- YES -> Use Mistral SDK directly
| Examples: Codestral FIM, Mistral Agents,
| Voxtral audio, OCR, custom endpoints
+-- NO -> Mistral SDK is simplest for Mistral-only use
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Using
response_format(snake_case) instead ofresponseFormat(camelCase) -- silently ignored, no error thrown - Using
input(singular) for embeddings instead ofinputs(plural) -- Mistral-specific naming - Not configuring retries for production (SDK defaults to
strategy: "none"-- zero retries) - Hardcoding API keys instead of using environment variables
- Accessing
chunk.choices[0]?.delta?.contentdirectly on streaming events instead ofevent.data.choices[0]?.delta?.content
Medium Priority Issues:
- Not setting
timeoutMsfor production (default is-1, meaning no timeout -- requests can hang indefinitely) - Using
max_tokensinstead ofmaxTokens(camelCase SDK convention) - Missing
systemrole message for behavior guidance - Using
tool_choiceinstead oftoolChoice - Using
tool_callsinstead oftoolCallswhen reading responses
Common Mistakes:
- Importing from
"mistralai"instead of"@mistralai/mistralai"-- the correct package name has the org scope - Using CommonJS
require()-- the package is ESM-only, useimportorawait import() - Confusing Mistral's
imageUrl: "url"(flat string) with OpenAI'simage_url: { url: "..." }(nested object) - Using
client.chat.completions.create()(OpenAI pattern) instead ofclient.chat.complete()(Mistral pattern) - Assuming embedding dimensions match OpenAI's --
mistral-embedreturns 1024-dimensional vectors, not 1536
Gotchas & Edge Cases:
- The SDK is ESM-only. In CommonJS projects, you must use
const { Mistral } = await import("@mistralai/mistralai"). - Streaming content may be
string | string[]-- cast or check type when writing to stdout. chat.parse()requires a Zod schema passed toresponseFormat-- it does not accept{ type: "json_object" }.- The
apiKeyconstructor option accepts a string OR an async function() => Promise<string>for dynamic key rotation. - Model aliases like
mistral-large-latestresolve to the latest version of that model tier. Pin to specific versions (e.g.,mistral-large-3-25-12) for reproducibility. toolChoice: "any"forces the model to call a tool.toolChoice: "auto"lets the model decide.toolChoice: "none"prevents tool calls.parallelToolCalls: falseforces sequential tool calling (defaulttrueallows parallel).- FIM endpoint (
fim.complete()) usesprompt+suffixparameters, NOT themessagesarray. safePrompt: trueinjects Mistral's safety system prompt before your messages.- The SDK provides standalone functions (e.g.,
chatComplete()from"@mistralai/mistralai/funcs/chatComplete.js") for tree-shaking in browser/edge runtimes.
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST use responseFormat (camelCase) in SDK calls -- NOT response_format (snake_case). The SDK uses camelCase property names throughout.)
(You MUST configure retries explicitly -- the SDK defaults to strategy: "none" (no retries), unlike OpenAI's SDK which retries automatically)
(You MUST consume streaming results with for await (const event of result) and access content via event.data.choices[0]?.delta?.content -- the event shape differs from OpenAI)
(You MUST never hardcode API keys -- use process.env["MISTRAL_API_KEY"] with the bracket notation the SDK documents)
(You MUST use client.chat.parse() with a Zod schema for structured outputs -- NOT manual JSON.parse() on completion content)
Failure to follow these rules will produce broken API calls (snake_case properties silently ignored), unreliable production services (no retries), or incorrectly parsed streaming data.
</critical_reminders>
Files (skills)
-
examples
-
chat.md 5.7 KB
# Mistral SDK -- Chat & Streaming Examples > Chat completions, streaming with async iteration, multi-turn conversations, and token tracking. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [function-calling.md](function-calling.md) -- Tool/function calling - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision - [codestral.md](codestral.md) -- Codestral FIM code completion --- ## Basic Chat Completion ```typescript // basic-chat.ts import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); async function chat(userMessage: string): Promise<string> { const result = await client.chat.complete({ model: "mistral-large-latest", messages: [ { role: "system", content: "You are a helpful assistant. Be concise." }, { role: "user", content: userMessage }, ], }); const content = result?.choices?.[0]?.message?.content; if (!content) { throw new Error("No content in response"); } // Log token usage if (result.usage) { console.log(`Tokens: ${result.usage.totalTokens}`); } return typeof content === "string" ? content : content.join(""); } const answer = await chat("What is TypeScript in one sentence?"); console.log(answer); ``` **Why good:** Safe optional chaining on nullable fields, handles `content` being `string | string[]`, token tracking --- ## Streaming with Async Iteration ```typescript // streaming-chat.ts import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); const result = await client.chat.stream({ model: "mistral-large-latest", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: "Explain async/await in TypeScript." }, ], }); // IMPORTANT: events wrap data in event.data -- not directly on event for await (const event of result) { const content = event.data.choices[0]?.delta?.content; if (content) { // content may be string | string[] depending on model process.stdout.write( typeof content === "string" ? content : content.join(""), ); } } console.log(); // newline ``` **Why good:** Accesses `event.data.choices[0]` (not `event.choices[0]`), handles string union type ### BAD: OpenAI-style streaming (wrong for Mistral) ```typescript // BAD: Trying OpenAI's event shape for await (const chunk of result) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); // WRONG: no .data wrapper } ``` **Why bad:** Mistral wraps streaming data in `event.data` -- accessing `chunk.choices` directly fails silently or throws --- ## Multi-Turn Conversation ```typescript import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); interface ChatMessage { role: "system" | "user" | "assistant"; content: string; } const messages: ChatMessage[] = [ { role: "system", content: "You are a TypeScript expert." }, { role: "user", content: "What is a union type?" }, ]; const first = await client.chat.complete({ model: "mistral-large-latest", messages, }); const firstContent = first?.choices?.[0]?.message?.content; if (firstContent) { // Append assistant response for next turn messages.push({ role: "assistant", content: typeof firstContent === "string" ? firstContent : firstContent.join(""), }); messages.push({ role: "user", content: "Give me a real-world example." }); const followUp = await client.chat.complete({ model: "mistral-large-latest", messages, }); console.log(followUp?.choices?.[0]?.message?.content); } ``` --- ## Controlling Output Length and Temperature ```typescript const MAX_TOKENS = 500; const result = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Summarize this article." }], maxTokens: MAX_TOKENS, temperature: 0, // Deterministic output }); const finishReason = result?.choices?.[0]?.finishReason; if (finishReason === "length") { console.warn("Output was truncated -- increase maxTokens"); } ``` **Why good:** Uses `maxTokens` (camelCase), checks `finishReason` (camelCase) for truncation --- ## Token Usage Tracking ```typescript const result = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Hello" }], }); const usage = result?.usage; if (usage) { console.log(`Prompt tokens: ${usage.promptTokens}`); console.log(`Completion tokens: ${usage.completionTokens}`); console.log(`Total tokens: ${usage.totalTokens}`); } ``` **Why good:** Uses camelCase properties (`promptTokens`, not `prompt_tokens`) --- ## Streaming with Abort ```typescript const controller = new AbortController(); const ABORT_TIMEOUT_MS = 5_000; setTimeout(() => controller.abort(), ABORT_TIMEOUT_MS); try { const result = await client.chat.stream( { model: "mistral-large-latest", messages: [{ role: "user", content: "Tell me a long story." }], }, { fetchOptions: { signal: controller.signal }, }, ); for await (const event of result) { const content = event.data.choices[0]?.delta?.content; if (content) { process.stdout.write( typeof content === "string" ? content : content.join(""), ); } } } catch (error) { if (error instanceof Error && error.name === "AbortError") { console.log("\nStream aborted"); } else { throw error; } } ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
codestral.md 4 KB
# Mistral SDK -- Codestral FIM Examples > Fill-in-middle code completion and code generation with Codestral. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat completions and streaming - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [function-calling.md](function-calling.md) -- Tool/function calling - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision --- ## Basic Fill-in-Middle (FIM) ```typescript // fim-completion.ts import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); // FIM: provide code before (prompt) and after (suffix) the cursor const result = await client.fim.complete({ model: "codestral-latest", prompt: "function fibonacci(n: number): number {\n if (n <= 1) return n;\n", suffix: "}\n\nconsole.log(fibonacci(10));", temperature: 0, }); const completion = result.choices?.[0]?.message?.content; console.log("Generated code:", completion); // The model fills in the gap: " return fibonacci(n - 1) + fibonacci(n - 2);\n" ``` **Why good:** Uses dedicated `fim.complete()` endpoint (not chat), separate `prompt` + `suffix`, deterministic with `temperature: 0` --- ## FIM with Stop Sequences ```typescript const result = await client.fim.complete({ model: "codestral-latest", prompt: "class UserService {\n private users: User[] = [];\n\n ", suffix: "\n\n getUsers(): User[] {\n return this.users;\n }\n}", temperature: 0, stop: ["\n\n"], // Stop at double newline to prevent over-generation }); const code = result.choices?.[0]?.message?.content; console.log(code); ``` --- ## FIM vs Chat for Code Generation Use FIM when you have surrounding code context. Use chat when you need full function generation from a description. ```typescript // FIM: filling a gap in existing code const fimResult = await client.fim.complete({ model: "codestral-latest", prompt: "const sorted = array.", suffix: ";\nconsole.log(sorted);", temperature: 0, }); // Model completes: "sort((a, b) => a - b)" // Chat: generating from a description const chatResult = await client.chat.complete({ model: "codestral-latest", messages: [ { role: "system", content: "You are a TypeScript code generator. Return only code, no explanations.", }, { role: "user", content: "Write a function that sorts an array of numbers in ascending order.", }, ], temperature: 0, }); ``` **When to use FIM:** IDE autocomplete, code insertion at cursor position, completing partial code **When to use Chat:** Generating entire functions, explaining code, code review --- ## Code Completion with Max Tokens ```typescript const MAX_COMPLETION_TOKENS = 200; const result = await client.fim.complete({ model: "codestral-latest", prompt: "async function fetchUsers(): Promise<User[]> {\n ", suffix: "\n}", temperature: 0, maxTokens: MAX_COMPLETION_TOKENS, }); const finishReason = result.choices?.[0]?.finishReason; if (finishReason === "length") { console.warn("Code completion was truncated -- increase maxTokens"); } ``` --- ## Code Embeddings with Codestral Embed For code-specific semantic search, use `codestral-embed-latest` instead of `mistral-embed`. ```typescript const CODE_EMBEDDING_MODEL = "codestral-embed-latest"; const result = await client.embeddings.create({ model: CODE_EMBEDDING_MODEL, inputs: [ "function add(a: number, b: number): number { return a + b; }", "const sum = (x: number, y: number): number => x + y;", "class Calculator { add(a: number, b: number) { return a + b; } }", ], }); // Use for code search, duplicate detection, or semantic code navigation const vectors = result.data?.map((item) => item.embedding) ?? []; console.log(`Generated ${vectors.length} code embeddings`); ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
core.md 6.4 KB
# Mistral SDK -- Setup & Configuration Examples > Client initialization, environment config, production settings, error handling, custom HTTP client, and retry configuration. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [chat.md](chat.md) -- Chat completions and streaming - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [function-calling.md](function-calling.md) -- Tool/function calling - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision - [codestral.md](codestral.md) -- Codestral FIM code completion --- ## Basic Client Setup ```typescript // lib/mistral.ts import { Mistral } from "@mistralai/mistralai"; // Reads MISTRAL_API_KEY from env const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "", }); export { client }; ``` --- ## Production Configuration ```typescript // lib/mistral.ts import { Mistral } from "@mistralai/mistralai"; const TIMEOUT_MS = 30_000; const INITIAL_RETRY_INTERVAL_MS = 1_000; const MAX_RETRY_INTERVAL_MS = 30_000; const RETRY_EXPONENT = 1.5; const MAX_ELAPSED_TIME_MS = 120_000; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "", timeoutMs: TIMEOUT_MS, retryConfig: { strategy: "backoff", backoff: { initialInterval: INITIAL_RETRY_INTERVAL_MS, maxInterval: MAX_RETRY_INTERVAL_MS, exponent: RETRY_EXPONENT, maxElapsedTime: MAX_ELAPSED_TIME_MS, }, retryConnectionErrors: true, }, }); export { client }; ``` **Why good:** Explicit retry config (SDK defaults to zero retries), named constants, sensible backoff curve --- ## Async API Key Provider ```typescript // lib/mistral.ts -- dynamic key rotation import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: async () => { // Fetch from secrets manager, vault, or key rotation service return await getSecureApiKey(); }, timeoutMs: 30_000, }); export { client }; ``` **Why good:** Supports key rotation without restarting the process -- the SDK calls the function before each request --- ## Custom HTTP Client with Hooks ```typescript // lib/mistral.ts -- custom HTTP client import { Mistral } from "@mistralai/mistralai"; import { HTTPClient } from "@mistralai/mistralai/lib/http"; const REQUEST_TIMEOUT_MS = 10_000; const httpClient = new HTTPClient({ fetcher: (request) => fetch(request), }); // Add timeout to every request httpClient.addHook("beforeRequest", (request) => { const nextRequest = new Request(request, { signal: request.signal || AbortSignal.timeout(REQUEST_TIMEOUT_MS), }); nextRequest.headers.set("x-custom-header", "custom-value"); return nextRequest; }); // Log errors httpClient.addHook("requestError", (error, request) => { console.error( `Mistral request failed: ${request.method} ${request.url}`, error, ); }); const client = new Mistral({ httpClient, apiKey: process.env["MISTRAL_API_KEY"] ?? "", }); export { client }; ``` --- ## Production Error Handling ```typescript // error-handling.ts import { Mistral } from "@mistralai/mistralai"; import { SDKError, SDKValidationError, HTTPValidationError, } from "@mistralai/mistralai/models/errors"; const TIMEOUT_MS = 30_000; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "", timeoutMs: TIMEOUT_MS, retryConfig: { strategy: "backoff", backoff: { initialInterval: 1_000, maxInterval: 30_000, exponent: 1.5, maxElapsedTime: 60_000, }, retryConnectionErrors: true, }, }); async function safeCompletion(prompt: string): Promise<string | null> { try { const result = await client.chat.complete({ model: "mistral-large-latest", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: prompt }, ], }); const content = result?.choices?.[0]?.message?.content; if (!content) { console.warn("No content in response"); return null; } // Log token usage for cost tracking if (result.usage) { console.log( `Tokens: prompt=${result.usage.promptTokens}, completion=${result.usage.completionTokens}, total=${result.usage.totalTokens}`, ); } return typeof content === "string" ? content : content.join(""); } catch (error) { if (error instanceof HTTPValidationError) { // 422 -- invalid request shape console.error("Validation error:", error.message); return null; } if (error instanceof SDKValidationError) { // Client-side input validation failure console.error("Input validation error:", error.message); return null; } if (error instanceof SDKError) { // General API error (4xx, 5xx) console.error(`API error [${error.statusCode}]: ${error.message}`); if (error.statusCode === 401) { throw new Error( "Invalid API key. Check MISTRAL_API_KEY environment variable.", ); } if (error.statusCode === 429) { console.error("Rate limited. Retries exhausted."); return null; } return null; } // Unknown errors should be re-thrown throw error; } } const result = await safeCompletion("Hello!"); if (result) { console.log(result); } else { console.error("Failed to get completion"); } ``` --- ## Per-Request Override ```typescript // Override retries and timeout for a single request const result = await client.chat.complete( { model: "mistral-large-latest", messages: [{ role: "user", content: "Quick question" }], }, { retries: { strategy: "backoff", backoff: { initialInterval: 500, maxInterval: 5_000, exponent: 2, maxElapsedTime: 15_000, }, }, fetchOptions: { signal: AbortSignal.timeout(10_000), }, }, ); ``` --- ## Debug Logging ```typescript // Enable via environment variable // MISTRAL_DEBUG=true node app.js // Or via constructor const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "", debugLogger: console, }); // Or with a custom logger const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "", debugLogger: { log: (msg: string) => logger.info(msg), error: (msg: string) => logger.error(msg), warn: (msg: string) => logger.warn(msg), debug: (msg: string) => logger.debug(msg), }, }); ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
embeddings-vision.md 6 KB
# Mistral SDK -- Embeddings & Vision Examples > Embeddings for semantic search and vision/image inputs with vision-capable models. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat completions and streaming - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [function-calling.md](function-calling.md) -- Tool/function calling - [codestral.md](codestral.md) -- Codestral FIM code completion --- ## Embeddings and Semantic Search ```typescript // embeddings.ts import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); const EMBEDDING_MODEL = "mistral-embed"; const SIMILARITY_THRESHOLD = 0.7; const TOP_K = 3; function cosineSimilarity(a: number[], b: number[]): number { let dot = 0; let normA = 0; let normB = 0; for (let i = 0; i < a.length; i++) { dot += a[i] * b[i]; normA += a[i] * a[i]; normB += b[i] * b[i]; } return dot / (Math.sqrt(normA) * Math.sqrt(normB)); } // Index documents -- batch multiple inputs in one call const documents = [ "TypeScript provides static type checking for JavaScript.", "React is a library for building user interfaces.", "Node.js is a JavaScript runtime built on V8.", "PostgreSQL is a powerful relational database.", "Docker containers package applications with dependencies.", ]; const docEmbeddings = await client.embeddings.create({ model: EMBEDDING_MODEL, inputs: documents, // NOTE: 'inputs' (plural), not 'input' }); const indexedDocs = documents.map((text, i) => ({ text, embedding: docEmbeddings.data?.[i]?.embedding ?? [], })); // Search async function search( query: string, ): Promise<Array<{ text: string; score: number }>> { const queryEmbedding = await client.embeddings.create({ model: EMBEDDING_MODEL, inputs: [query], }); const queryVector = queryEmbedding.data?.[0]?.embedding ?? []; return indexedDocs .map((doc) => ({ text: doc.text, score: cosineSimilarity(queryVector, doc.embedding), })) .filter((r) => r.score > SIMILARITY_THRESHOLD) .sort((a, b) => b.score - a.score) .slice(0, TOP_K); } const results = await search("What is TypeScript?"); results.forEach((r) => { console.log(`[${r.score.toFixed(3)}] ${r.text}`); }); ``` **Why good:** Uses `inputs` (plural, Mistral-specific), batches documents in one call, named constants --- ## Embedding Gotchas ```typescript // GOOD: Mistral uses 'inputs' (plural) const result = await client.embeddings.create({ model: "mistral-embed", inputs: ["text to embed"], }); // Returns 1024-dimensional vectors // BAD: OpenAI uses 'input' (singular), dimensions differ const result = await client.embeddings.create({ model: "mistral-embed", input: ["text to embed"], // WRONG: use 'inputs' }); // OpenAI returns 1536 dimensions, Mistral returns 1024 ``` **Why bad:** Wrong parameter name, wrong dimension expectations when migrating from OpenAI --- ## Vision -- Image from URL ```typescript // vision.ts import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); async function analyzeImageUrl( imageUrl: string, question: string, ): Promise<string> { const result = await client.chat.complete({ model: "mistral-small-latest", // Vision-capable model messages: [ { role: "user", content: [ { type: "text", text: question }, { type: "image_url", imageUrl: imageUrl, // NOTE: flat string, not { url: "..." } }, ], }, ], }); const content = result?.choices?.[0]?.message?.content; return typeof content === "string" ? content : (content?.join("") ?? ""); } const description = await analyzeImageUrl( "https://example.com/photo.jpg", "Describe what you see in this image.", ); console.log(description); ``` **Why good:** Uses `imageUrl` (camelCase, flat string) -- NOT OpenAI's `image_url: { url: "..." }` nested object --- ## Vision -- Local Image (Base64) ```typescript import { Mistral } from "@mistralai/mistralai"; import { readFileSync } from "node:fs"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); async function analyzeLocalImage( imagePath: string, question: string, ): Promise<string> { const imageBuffer = readFileSync(imagePath); const base64Image = imageBuffer.toString("base64"); const mimeType = imagePath.endsWith(".png") ? "image/png" : "image/jpeg"; const result = await client.chat.complete({ model: "mistral-small-latest", messages: [ { role: "user", content: [ { type: "text", text: question }, { type: "image_url", imageUrl: `data:${mimeType};base64,${base64Image}`, }, ], }, ], }); const content = result?.choices?.[0]?.message?.content; return typeof content === "string" ? content : (content?.join("") ?? ""); } ``` --- ## Vision -- Multiple Images ```typescript const result = await client.chat.complete({ model: "mistral-small-latest", messages: [ { role: "user", content: [ { type: "text", text: "Compare these two images." }, { type: "image_url", imageUrl: "https://example.com/image1.jpg", }, { type: "image_url", imageUrl: "https://example.com/image2.jpg", }, ], }, ], }); ``` --- ## Vision Model Selection ``` Vision-capable models (as of 2026): - mistral-small-latest (Mistral Small 4) - mistral-medium-latest (Mistral Medium 3.1) - mistral-large-latest (Mistral Large 3) - ministral-14b-latest, ministral-8b-latest, ministral-3b-latest (Ministral 3) ``` Most current Mistral models support vision. Use `mistral-small-latest` for cost-efficient image analysis. --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
function-calling.md 6.8 KB
# Mistral SDK -- Function Calling Examples > Tool definitions, tool call loop, parallel tool calls, streaming function calling. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat completions and streaming - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision - [codestral.md](codestral.md) -- Codestral FIM code completion --- ## Basic Function Calling ```typescript // function-calling.ts import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); const tools = [ { type: "function" as const, function: { name: "get_weather", description: "Get the current weather for a location", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, unit: { type: "string", enum: ["celsius", "fahrenheit"], description: "Temperature unit", }, }, required: ["location"], }, }, }, ]; const result = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "What is the weather in Paris?" }], tools, toolChoice: "any", // Forces tool use }); const toolCall = result?.choices?.[0]?.message?.toolCalls?.[0]; if (toolCall) { const args = JSON.parse(toolCall.function.arguments); console.log(`Call ${toolCall.function.name} with:`, args); console.log(`Tool call ID: ${toolCall.id}`); } ``` **Why good:** Uses `toolChoice` (camelCase), `toolCalls` (camelCase), `as const` for type literal, checks for tool calls --- ## Complete Tool Call Loop The standard pattern: send message with tools -> get tool calls -> execute -> send results back. ```typescript // tool-loop.ts import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); // Tool implementations async function getWeather(args: { location: string; unit?: string; }): Promise<string> { return JSON.stringify({ location: args.location, temperature: 22, unit: args.unit ?? "celsius", condition: "sunny", }); } async function searchDatabase(args: { query: string }): Promise<string> { return JSON.stringify({ results: [{ id: 1, title: `Result for: ${args.query}` }], }); } const toolImplementations: Record< string, (args: Record<string, unknown>) => Promise<string> > = { get_weather: getWeather as (args: Record<string, unknown>) => Promise<string>, search_database: searchDatabase as ( args: Record<string, unknown>, ) => Promise<string>, }; const tools = [ { type: "function" as const, function: { name: "get_weather", description: "Get current weather for a location", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, unit: { type: "string", enum: ["celsius", "fahrenheit"] }, }, required: ["location"], }, }, }, { type: "function" as const, function: { name: "search_database", description: "Search the knowledge base", parameters: { type: "object", properties: { query: { type: "string", description: "Search query" }, }, required: ["query"], }, }, }, ]; interface ChatMessage { role: "system" | "user" | "assistant" | "tool"; content: string; name?: string; toolCallId?: string; toolCalls?: Array<{ id: string; function: { name: string; arguments: string }; }>; } const MAX_TOOL_ITERATIONS = 5; const messages: ChatMessage[] = [ { role: "system", content: "You help users with weather and search queries.", }, { role: "user", content: "What is the weather in London and search for TypeScript guides?", }, ]; let iterations = 0; while (iterations < MAX_TOOL_ITERATIONS) { const result = await client.chat.complete({ model: "mistral-large-latest", messages, tools, }); const choice = result?.choices?.[0]; if (!choice) break; const assistantMessage = choice.message; // Check if the model wants to call tools if (!assistantMessage?.toolCalls || assistantMessage.toolCalls.length === 0) { // No tool calls -- model has a final answer console.log("Final answer:", assistantMessage?.content); break; } // Add assistant message with tool calls to history messages.push({ role: "assistant", content: typeof assistantMessage.content === "string" ? assistantMessage.content : "", toolCalls: assistantMessage.toolCalls.map((tc) => ({ id: tc.id ?? "", function: { name: tc.function.name, arguments: tc.function.arguments }, })), }); // Execute each tool call and add results for (const toolCall of assistantMessage.toolCalls) { const fnName = toolCall.function.name; const args = JSON.parse(toolCall.function.arguments); console.log(`Calling ${fnName}(${JSON.stringify(args)})`); const impl = toolImplementations[fnName]; if (!impl) { throw new Error(`Unknown tool: ${fnName}`); } const toolResult = await impl(args); messages.push({ role: "tool", name: fnName, content: toolResult, toolCallId: toolCall.id ?? "", }); } iterations++; } ``` **Why good:** Bounded loop with MAX_TOOL_ITERATIONS, proper tool message format with `toolCallId`, handles parallel tool calls --- ## Tool Choice Options ```typescript // "auto" (default) -- model decides whether to call tools const result1 = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Hello" }], tools, toolChoice: "auto", }); // "any" -- forces the model to call at least one tool const result2 = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Get weather in Tokyo" }], tools, toolChoice: "any", }); // "none" -- prevents any tool calls const result3 = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Tell me about weather" }], tools, toolChoice: "none", }); ``` --- ## Sequential Tool Calls ```typescript // Force sequential tool calling (disable parallel calls) const result = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Weather in Paris and London?" }], tools, parallelToolCalls: false, // Forces sequential -- model calls one tool at a time }); ``` **Why good:** `parallelToolCalls: false` ensures tools are called one at a time, useful for dependent operations --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
structured-output.md 5.3 KB
# Mistral SDK -- Structured Output Examples > Type-safe structured responses with Zod via `chat.parse()`, and JSON mode via `responseFormat`. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat completions and streaming - [function-calling.md](function-calling.md) -- Tool/function calling - [embeddings-vision.md](embeddings-vision.md) -- Embeddings and vision - [codestral.md](codestral.md) -- Codestral FIM code completion --- ## Structured Output with `chat.parse()` and Zod ```typescript // structured-output.ts import { Mistral } from "@mistralai/mistralai"; import { z } from "zod"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); const ArticleSummary = z.object({ title: z.string(), summary: z.string(), keyPoints: z.array(z.string()), sentiment: z.enum(["positive", "negative", "neutral"]), }); type ArticleSummary = z.infer<typeof ArticleSummary>; const MAX_TOKENS = 512; async function extractSummary( articleText: string, ): Promise<ArticleSummary | null> { const result = await client.chat.parse({ model: "mistral-large-latest", messages: [ { role: "system", content: "Extract a structured summary from the article.", }, { role: "user", content: articleText }, ], responseFormat: ArticleSummary, // Pass Zod schema directly maxTokens: MAX_TOKENS, temperature: 0, }); // Access the typed parsed result const parsed = result.choices?.[0]?.message?.parsed; return parsed ?? null; } const article = ` TypeScript 5.8 brings improved type inference, better error messages, and performance optimizations focused on developer experience. `; const summary = await extractSummary(article); if (summary) { console.log(`Title: ${summary.title}`); console.log(`Sentiment: ${summary.sentiment}`); console.log("Key Points:"); summary.keyPoints.forEach((point) => console.log(` - ${point}`)); } ``` **Why good:** Zod schema passed directly to `responseFormat`, `message.parsed` is fully typed, named constants --- ## JSON Mode (Simple) When you need JSON output without strict schema validation, use `responseFormat: { type: "json_object" }`. ```typescript import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); const result = await client.chat.complete({ model: "mistral-large-latest", messages: [ { role: "user", content: "List the top 3 programming languages with their use cases. Return as JSON.", }, ], responseFormat: { type: "json_object" }, }); const content = result?.choices?.[0]?.message?.content; if (typeof content === "string") { const data = JSON.parse(content); console.log(data); } ``` **Why good:** Uses `responseFormat` (camelCase), explicitly asks for JSON in the prompt (recommended even with JSON mode) ### BAD: snake_case response_format ```typescript // BAD: Using REST API naming in the SDK const result = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Return JSON" }], response_format: { type: "json_object" }, // WRONG: silently ignored }); ``` **Why bad:** The SDK uses camelCase -- `response_format` is silently ignored, model returns plain text --- ## Complex Schema with Nested Objects ```typescript import { Mistral } from "@mistralai/mistralai"; import { z } from "zod"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); const Address = z.object({ street: z.string(), city: z.string(), country: z.string(), postalCode: z.string(), }); const Person = z.object({ name: z.string(), age: z.number(), email: z.string().email(), addresses: z.array(Address), occupation: z.string().optional(), }); const MAX_TOKENS = 512; const result = await client.chat.parse({ model: "mistral-large-latest", messages: [ { role: "system", content: "Extract person details from the text." }, { role: "user", content: "John Smith, 34, works as a software engineer. " + "Email: john@example.com. Lives at 123 Main St, Paris, France 75001.", }, ], responseFormat: Person, maxTokens: MAX_TOKENS, temperature: 0, }); const person = result.choices?.[0]?.message?.parsed; if (person) { console.log(`${person.name}, age ${person.age}`); console.log(`Email: ${person.email}`); person.addresses.forEach((addr) => { console.log(`Address: ${addr.street}, ${addr.city}, ${addr.country}`); }); } ``` --- ## Accessing Raw vs Parsed Content ```typescript const result = await client.chat.parse({ model: "mistral-large-latest", messages: [ { role: "system", content: "Extract book information." }, { role: "user", content: "I read 'Dune' by Frank Herbert." }, ], responseFormat: z.object({ name: z.string(), authors: z.array(z.string()), }), }); const message = result.choices?.[0]?.message; // Raw JSON string (always available) console.log("Raw:", message?.content); // Parsed and typed object (available via chat.parse) console.log("Parsed:", message?.parsed); ``` **Why good:** Shows both access paths -- `content` for raw JSON, `parsed` for typed object --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
-
-
reference.md 9.9 KB
# Mistral SDK Quick Reference > Client configuration, model IDs, API methods, error types, and SDK conventions. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Package Installation ```bash # Core package (ESM-only) npm install @mistralai/mistralai # For structured outputs (recommended) npm install zod ``` --- ## Client Configuration ```typescript import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "", // string or async () => Promise<string> timeoutMs: 30_000, // Request timeout in ms (default: -1 = no timeout) server: "eu", // Named server selection (default: "eu") serverURL: "https://custom-api.example.com", // Custom endpoint override retryConfig: { // Default: { strategy: "none" } -- NO RETRIES strategy: "backoff", backoff: { initialInterval: 1_000, maxInterval: 30_000, exponent: 1.5, maxElapsedTime: 120_000, }, retryConnectionErrors: true, }, debugLogger: console, // Enable debug logging (or set MISTRAL_DEBUG=true) }); ``` ### Environment Variables | Variable | Purpose | | ----------------- | ----------------------------- | | `MISTRAL_API_KEY` | API key (required) | | `MISTRAL_DEBUG` | Enable debug logging (`true`) | ### Configuration Priority ``` Request-level options (highest) -> Client initialization options -> Environment variables -> SDK defaults (lowest) ``` --- ## Model IDs ### Language Models (Chat / Text Generation) | Model ID | Alias | Use Case | Context | | -------------------------- | ----------------------- | ------------------------------ | ------- | | `mistral-large-3-25-12` | `mistral-large-latest` | Most capable, general purpose | 256K | | `mistral-medium-3-1-25-08` | `mistral-medium-latest` | Balanced cost/performance | 256K | | `mistral-small-4-0-26-03` | `mistral-small-latest` | Cost-efficient, vision-capable | 256K | | `ministral-3-14b-25-12` | `ministral-14b-latest` | Compact, multimodal | 128K | | `ministral-3-8b-25-12` | `ministral-8b-latest` | Compact, multimodal | 128K | | `ministral-3-3b-25-12` | `ministral-3b-latest` | Edge / minimal | 128K | ### Reasoning Models | Model ID | Alias | Use Case | | ---------------------------- | ------------------------- | ----------------- | | `magistral-medium-1-2-25-09` | `magistral-medium-latest` | Complex reasoning | | `magistral-small-1-2-25-09` | `magistral-small-latest` | Fast reasoning | ### Code Models | Model ID | Alias | Use Case | | ------------------ | ------------------ | ----------------------------- | | `codestral-25-08` | `codestral-latest` | Code generation + FIM | | `devstral-2-25-12` | `devstral-latest` | Code agents, codebase explore | ### Embedding Models | Model ID | Alias | Dimensions | Max Input | | ----------------------- | ------------------------ | ---------- | ----------- | | `mistral-embed-23-12` | `mistral-embed` | 1024 | 8192 tokens | | `codestral-embed-25-05` | `codestral-embed-latest` | 1024 | 8192 tokens | ### Specialist Models | Model ID | Use Case | | ------------------------------- | ------------------- | | `mistral-ocr-3-25-12` | Document OCR | | `mistral-moderation-26-03` | Content moderation | | `voxtral-mini-transcribe-26-02` | Audio transcription | --- ## API Methods Reference ### Chat Completions ```typescript // Standard completion const result = await client.chat.complete({ model: "mistral-large-latest", // Required messages: [], // Required: ChatCompletionRequestMessage[] temperature: 0.7, // 0.0-1.5 (default: 0.7) maxTokens: 1000, // Max output tokens topP: 1, // Nucleus sampling (default: 1) frequencyPenalty: 0, // Penalize frequent tokens presencePenalty: 0, // Penalize present tokens tools: [], // Tool definitions toolChoice: "auto", // "auto" | "any" | "none" | "required" parallelToolCalls: true, // Allow parallel tool calls responseFormat: { type: "text" }, // "text" | "json_object" safePrompt: false, // Inject safety prompt stop: undefined, // string | string[] -- stop sequences }); // Structured output parsing with Zod const result = await client.chat.parse({ model: "mistral-large-latest", messages: [], responseFormat: zodSchema, // Pass Zod schema directly maxTokens: 256, temperature: 0, }); // Streaming const result = await client.chat.stream({ model: "mistral-large-latest", messages: [], }); for await (const event of result) { // event.data.choices[0]?.delta?.content } ``` ### FIM (Fill-in-Middle) ```typescript const result = await client.fim.complete({ model: "codestral-latest", // Required prompt: "", // Code before cursor (required) suffix: "", // Code after cursor (optional) temperature: 0, // Sampling temperature maxTokens: 1000, // Max output tokens stop: undefined, // Stop sequences }); ``` ### Embeddings ```typescript const result = await client.embeddings.create({ model: "mistral-embed", // Required inputs: [], // string[] (NOTE: plural 'inputs', not 'input') }); ``` ### Agents ```typescript const result = await client.agents.complete({ agentId: "<id>", // Required: pre-configured agent ID messages: [], // Required responseFormat: { type: "text" }, }); // Streaming const result = await client.agents.stream({ agentId: "<id>", messages: [], }); ``` ### Files ```typescript import { openAsBlob } from "node:fs"; // Upload const result = await client.files.upload({ file: await openAsBlob("data.jsonl"), }); // List, retrieve, delete const files = await client.files.list(); const file = await client.files.retrieve({ fileId: "file-abc123" }); await client.files.delete({ fileId: "file-abc123" }); ``` ### Models ```typescript const models = await client.models.list(); const model = await client.models.retrieve({ modelId: "mistral-large-latest" }); ``` --- ## Error Types | Error Class | Description | | --------------------- | ------------------------------------------ | | `SDKError` | General API errors (4XX, 5XX status codes) | | `SDKValidationError` | Client-side input validation failure | | `HTTPValidationError` | Server-side validation (422) | | `ConnectionError` | Network connectivity issues | | `RequestTimeoutError` | Request timeout exceeded | | `RequestAbortedError` | Client-cancelled request | ```typescript import { SDKError, SDKValidationError, HTTPValidationError, } from "@mistralai/mistralai/models/errors"; ``` --- ## Retry Configuration ```typescript // Global retry (at client init) retryConfig: { strategy: "backoff", // "none" | "backoff" backoff: { initialInterval: 1_000, // ms between first retry maxInterval: 30_000, // ms max between retries exponent: 1.5, // backoff multiplier maxElapsedTime: 120_000, // ms total retry window }, retryConnectionErrors: true, } // Per-request retry override await client.chat.complete( { model: "mistral-large-latest", messages: [...] }, { retries: { strategy: "backoff", backoff: { initialInterval: 500, maxInterval: 5_000, exponent: 2, maxElapsedTime: 30_000 }, }, }, ); ``` --- ## Custom HTTP Client ```typescript import { HTTPClient } from "@mistralai/mistralai/lib/http"; const httpClient = new HTTPClient({ fetcher: (request) => fetch(request), }); httpClient.addHook("beforeRequest", (request) => { const nextRequest = new Request(request, { signal: request.signal || AbortSignal.timeout(5_000), }); nextRequest.headers.set("x-custom-header", "value"); return nextRequest; }); httpClient.addHook("requestError", (error, request) => { console.error(`Request failed: ${request.method} ${request.url}`, error); }); const client = new Mistral({ httpClient, apiKey: process.env["MISTRAL_API_KEY"] ?? "", }); ``` --- ## SDK vs REST API Property Names The SDK uses camelCase. The REST API uses snake_case. This table maps the most common properties: | SDK (camelCase) | REST API (snake_case) | | ------------------- | --------------------- | | `responseFormat` | `response_format` | | `maxTokens` | `max_tokens` | | `topP` | `top_p` | | `toolChoice` | `tool_choice` | | `toolCalls` | `tool_calls` | | `parallelToolCalls` | `parallel_tool_calls` | | `frequencyPenalty` | `frequency_penalty` | | `presencePenalty` | `presence_penalty` | | `safePrompt` | `safe_prompt` | | `imageUrl` | `image_url` | | `agentId` | `agent_id` | --- ## Standalone Functions (Tree-Shaking) For browser/edge runtimes, import standalone functions to reduce bundle size: ```typescript import { chatComplete } from "@mistralai/mistralai/funcs/chatComplete.js"; import { chatStream } from "@mistralai/mistralai/funcs/chatStream.js"; import { embeddingsCreate } from "@mistralai/mistralai/funcs/embeddingsCreate.js"; const res = await chatComplete(client, { model: "mistral-small-latest", messages: [{ role: "user", content: "Hello" }], }); if (res.ok) { console.log(res.value); } ``` --- ## Message Roles | Role | Description | | ----------- | ------------------------------------------------------ | | `system` | System instruction (behavior guidance) | | `user` | User input | | `assistant` | Model response | | `tool` | Tool result (requires `name`, `content`, `toolCallId`) | -
SKILL.md 21.5 KB
--- name: ai-provider-mistral-sdk description: Official Mistral AI TypeScript SDK patterns — client setup, chat completions, streaming, function calling, structured outputs, embeddings, vision, Codestral FIM, and production best practices --- # Mistral SDK Patterns > **Quick Guide:** Use `@mistralai/mistralai` (ESM-only) to interact with Mistral's API. Use `client.chat.complete()` for chat, `client.chat.stream()` for streaming (async iterable via `for await`), `client.chat.parse()` with a Zod schema for structured outputs, and `client.fim.complete()` for Codestral fill-in-middle code completion. The SDK uses `responseFormat` (camelCase) not `response_format`. Streaming events expose content via `event.data.choices[0]?.delta?.content`. Retries default to `strategy: "none"` -- you must configure them explicitly for production. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST use `responseFormat` (camelCase) in SDK calls -- NOT `response_format` (snake_case). The SDK uses camelCase property names throughout.)** **(You MUST configure retries explicitly -- the SDK defaults to `strategy: "none"` (no retries), unlike OpenAI's SDK which retries automatically)** **(You MUST consume streaming results with `for await (const event of result)` and access content via `event.data.choices[0]?.delta?.content` -- the event shape differs from OpenAI)** **(You MUST never hardcode API keys -- use `process.env["MISTRAL_API_KEY"]` with the bracket notation the SDK documents)** **(You MUST use `client.chat.parse()` with a Zod schema for structured outputs -- NOT manual `JSON.parse()` on completion content)** </critical_requirements> --- **Auto-detection:** Mistral, mistral, @mistralai/mistralai, client.chat.complete, client.chat.stream, client.chat.parse, client.fim.complete, client.embeddings.create, mistral-large, mistral-small, codestral, pixtral, ministral, magistral, devstral, MISTRAL_API_KEY, responseFormat, mistral-embed **When to use:** - Building applications that call Mistral models directly (Mistral Large, Small, Codestral, etc.) - Implementing chat completions with SSE streaming - Using Codestral for code generation and fill-in-middle (FIM) completion - Extracting structured data with `client.chat.parse()` and Zod schemas - Implementing function calling / tool use - Creating embeddings for RAG pipelines or semantic search - Processing images with vision-capable models (Mistral Small, Medium, Large, Ministral) - Using Mistral Agents API for pre-configured agent completions **Key patterns covered:** - Client initialization and configuration (retries, timeouts, custom HTTP client) - Chat completions (`chat.complete`) and streaming (`chat.stream`) - Structured outputs with `chat.parse()` and Zod schemas - Function calling / tool use with tool call loop - Embeddings (`embeddings.create`) with `mistral-embed` - Vision (image URL / base64 with vision-capable models) - Codestral FIM (`fim.complete`) for code completion - Error handling, retry configuration, and production patterns **When NOT to use:** - Multi-provider applications where you need to switch between Mistral, OpenAI, Anthropic, etc. -- use a unified provider SDK - React-specific chat UI hooks (`useChat`) -- use a framework-integrated AI SDK - When you need OpenAI-compatible endpoints -- use OpenAI SDK with Mistral's compatible endpoint instead --- ## Examples Index - [Core: Setup & Configuration](examples/core.md) -- Client init, production config, error handling, retries, custom HTTP client - [Chat & Streaming](examples/chat.md) -- Chat completions, streaming with async iteration, multi-turn - [Structured Output](examples/structured-output.md) -- `chat.parse()` with Zod, JSON mode, typed responses - [Function Calling](examples/function-calling.md) -- Tool definitions, tool call loop, streaming tools - [Embeddings & Vision](examples/embeddings-vision.md) -- Semantic search, image analysis with vision-capable models - [Codestral FIM](examples/codestral.md) -- Fill-in-middle code completion, code generation - [Quick API Reference](reference.md) -- Model IDs, method signatures, error types, configuration options --- <philosophy> ## Philosophy The `@mistralai/mistralai` SDK is **auto-generated from Mistral's OpenAPI spec using Speakeasy**, giving you a thin, type-safe wrapper over the REST API. It is ESM-only and uses camelCase property names (not snake_case like the REST API). **Core principles:** 1. **ESM-only** -- The package is published as ESM only. CommonJS projects must use `await import()`. This is a hard constraint, not optional. 2. **camelCase API surface** -- SDK properties use camelCase (`responseFormat`, `maxTokens`, `toolChoice`) even though the REST API uses snake_case. This catches OpenAI SDK migrants who write `response_format`. 3. **No automatic retries** -- Unlike OpenAI's SDK (2 retries by default), Mistral defaults to `strategy: "none"`. You must configure retries explicitly for production. 4. **Streaming via async iterables** -- `chat.stream()` returns an `EventStream` consumed with `for await...of`. Events have a `data` wrapper: `event.data.choices[0]?.delta?.content`. 5. **Structured outputs via `chat.parse()`** -- Pass a Zod schema directly to `responseFormat` and access `message.parsed` for typed results. No manual JSON schema construction needed. 6. **Codestral FIM** -- Dedicated `fim.complete()` endpoint for fill-in-middle code completion, separate from chat. **When to use the Mistral SDK directly:** - You only use Mistral models and want the simplest, most direct integration - You need Mistral-specific features (Codestral FIM, Mistral Agents, Voxtral audio) - You want minimal dependencies and zero abstraction overhead - You need the latest Mistral API features on day one **When NOT to use:** - You need to switch between providers (OpenAI, Anthropic, Mistral) -- use a unified provider SDK - You want React-specific chat UI hooks -- use a framework-integrated AI SDK - You want an OpenAI-compatible wrapper -- Mistral exposes an OpenAI-compatible endpoint, use the OpenAI SDK for that </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Setup Initialize the Mistral client. It reads `MISTRAL_API_KEY` from the environment. ```typescript // lib/mistral.ts -- basic setup import { Mistral } from "@mistralai/mistralai"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "", }); export { client }; ``` ```typescript // lib/mistral.ts -- production configuration import { Mistral } from "@mistralai/mistralai"; const TIMEOUT_MS = 30_000; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "", timeoutMs: TIMEOUT_MS, retryConfig: { strategy: "backoff", backoff: { initialInterval: 1_000, maxInterval: 30_000, exponent: 1.5, maxElapsedTime: 120_000, }, retryConnectionErrors: true, }, }); export { client }; ``` **Why good:** Explicit retry config (SDK defaults to no retries), named constants, env var with bracket notation **See:** [examples/core.md](examples/core.md) for custom HTTP client, async API key provider, error handling --- ### Pattern 2: Chat Completions Basic chat using `chat.complete()`. ```typescript const result = await client.chat.complete({ model: "mistral-large-latest", messages: [ { role: "system", content: "You are a helpful coding assistant." }, { role: "user", content: "Explain TypeScript generics." }, ], }); const content = result?.choices?.[0]?.message?.content; console.log(content); ``` **Why good:** Uses `system` role for instructions, safe optional chaining on nullable response ```typescript // BAD: Using snake_case properties (REST API style, not SDK style) const result = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "hello" }], response_format: { type: "json_object" }, // WRONG: use responseFormat max_tokens: 100, // WRONG: use maxTokens }); ``` **Why bad:** SDK uses camelCase properties -- `response_format` and `max_tokens` will be silently ignored **See:** [examples/chat.md](examples/chat.md) for multi-turn, token tracking, temperature control --- ### Pattern 3: Streaming Use `chat.stream()` for streaming. Events are async iterables. ```typescript const result = await client.chat.stream({ model: "mistral-large-latest", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: "Explain async/await in TypeScript." }, ], }); for await (const event of result) { const content = event.data.choices[0]?.delta?.content; if (content) { process.stdout.write(content as string); } } console.log(); ``` **Why good:** Proper `for await` iteration, accesses `event.data` (not `event` directly), handles nullable delta ```typescript // BAD: Trying to access content directly on event (OpenAI pattern) for await (const chunk of result) { process.stdout.write(chunk.choices[0]?.delta?.content ?? ""); // WRONG } ``` **Why bad:** Mistral streaming events wrap data in `event.data` -- direct access on `chunk` will fail **See:** [examples/chat.md](examples/chat.md) for complete streaming examples --- ### Pattern 4: Structured Outputs with Zod Use `chat.parse()` with a Zod schema for type-safe structured responses. ```typescript import { Mistral } from "@mistralai/mistralai"; import { z } from "zod"; const client = new Mistral({ apiKey: process.env["MISTRAL_API_KEY"] ?? "" }); const BookSchema = z.object({ name: z.string(), authors: z.array(z.string()), }); const MAX_TOKENS = 256; const result = await client.chat.parse({ model: "mistral-large-latest", messages: [ { role: "system", content: "Extract the book information." }, { role: "user", content: "I recently read 'Dune' by Frank Herbert." }, ], responseFormat: BookSchema, maxTokens: MAX_TOKENS, temperature: 0, }); const parsed = result.choices?.[0]?.message?.parsed; // parsed is typed as { name: string; authors: string[] } ``` **Why good:** Schema passed directly to `responseFormat`, `message.parsed` is fully typed, named constants **See:** [examples/structured-output.md](examples/structured-output.md) for JSON mode, complex schemas --- ### Pattern 5: Function Calling / Tool Use Define tools and handle the tool call loop. ```typescript const tools = [ { type: "function" as const, function: { name: "get_weather", description: "Get current weather for a city", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, }, required: ["location"], }, }, }, ]; const result = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Weather in Paris?" }], tools, toolChoice: "any", }); const toolCall = result?.choices?.[0]?.message?.toolCalls?.[0]; if (toolCall) { const args = JSON.parse(toolCall.function.arguments); console.log(`Call ${toolCall.function.name} with:`, args); } ``` **Why good:** Uses `toolChoice` (camelCase), `toolCalls` (camelCase), proper `as const` for type literal **See:** [examples/function-calling.md](examples/function-calling.md) for complete tool loop, parallel calls --- ### Pattern 6: Embeddings Create embeddings with `mistral-embed`. Note: uses `inputs` (plural), not `input`. ```typescript const EMBEDDING_MODEL = "mistral-embed"; const result = await client.embeddings.create({ model: EMBEDDING_MODEL, inputs: ["First document", "Second document", "Third document"], }); const vectors = result.data?.map((item) => item.embedding) ?? []; ``` **Why good:** Uses `inputs` (Mistral-specific, plural), named model constant, safe optional chaining ```typescript // BAD: Using singular 'input' (OpenAI pattern) const result = await client.embeddings.create({ model: "mistral-embed", input: ["First document"], // WRONG: Mistral uses 'inputs' (plural) }); ``` **Why bad:** Mistral SDK uses `inputs` (plural) -- `input` (singular) will error or be silently ignored **See:** [examples/embeddings-vision.md](examples/embeddings-vision.md) for cosine similarity, semantic search --- ### Pattern 7: Vision Send images to vision-capable models using multi-part content arrays. ```typescript const result = await client.chat.complete({ model: "mistral-small-latest", messages: [ { role: "user", content: [ { type: "text", text: "What is in this image?" }, { type: "image_url", imageUrl: "https://example.com/photo.jpg", }, ], }, ], }); ``` **Why good:** Uses `imageUrl` (camelCase string), not `image_url: { url }` (OpenAI's nested object pattern) **See:** [examples/embeddings-vision.md](examples/embeddings-vision.md) for base64 images, multiple images --- ### Pattern 8: Codestral FIM Fill-in-middle code completion using the dedicated FIM endpoint. ```typescript const result = await client.fim.complete({ model: "codestral-latest", prompt: "function fibonacci(n: number): number {\n if (n <= 1) return n;\n", suffix: "}\n\nconsole.log(fibonacci(10));", temperature: 0, }); const completion = result.choices?.[0]?.message?.content; // completion fills the gap between prompt and suffix ``` **Why good:** Dedicated FIM endpoint, separate `prompt` + `suffix` (not messages), deterministic with `temperature: 0` **See:** [examples/codestral.md](examples/codestral.md) for code generation patterns --- ### Pattern 9: Error Handling The SDK provides specific error types. Configure retries since the default is no retries. ```typescript import { Mistral } from "@mistralai/mistralai"; import { SDKError, SDKValidationError, HTTPValidationError, } from "@mistralai/mistralai/models/errors"; try { const result = await client.chat.complete({ model: "mistral-large-latest", messages: [{ role: "user", content: "Hello" }], }); } catch (error) { if (error instanceof HTTPValidationError) { console.error("Validation error:", error.message); } else if (error instanceof SDKValidationError) { console.error("Input validation error:", error.message); } else if (error instanceof SDKError) { console.error(`API error [${error.statusCode}]: ${error.message}`); } else { throw error; } } ``` **Why good:** Specific error types checked in order of specificity, re-throws unexpected errors **See:** [examples/core.md](examples/core.md) for full error handling, timeout handling, retry configuration </patterns> --- <performance> ## Performance Optimization ### Model Selection ``` General purpose (most capable) -> mistral-large-latest (Mistral Large 3) Balanced cost/quality -> mistral-medium-latest (Mistral Medium 3.1) Cost-sensitive / fast -> mistral-small-latest (Mistral Small 4) Edge / minimal -> ministral-3b-latest or ministral-8b-latest Complex reasoning -> magistral-medium-latest Code generation -> codestral-latest or devstral-latest Code completion (FIM) -> codestral-latest (dedicated FIM endpoint) Vision / images -> mistral-small-latest or mistral-large-latest Embeddings -> mistral-embed (1024 dimensions) Code embeddings -> codestral-embed-latest ``` ### Key Optimization Patterns - **Configure retries** -- Default is no retries. Always set `retryConfig` for production. - **Set timeouts** -- Default is no timeout (`-1`). Set `timeoutMs` to avoid hanging requests. - **Use `temperature: 0`** for deterministic output (enables server-side caching). - **Batch embedding inputs** -- Pass multiple strings to `inputs` array in one call. - **Use FIM for code completion** -- `fim.complete()` is purpose-built and more efficient than chat for code completion tasks. </performance> --- <decision_framework> ## Decision Framework ### Which Method to Use ``` What do you need? +-- Chat completion (text in, text out)? | +-- Need streaming? -> client.chat.stream() | +-- Need structured JSON? -> client.chat.parse() with Zod schema | +-- Basic completion? -> client.chat.complete() +-- Code completion / fill-in-middle? | +-- YES -> client.fim.complete() with Codestral +-- Embeddings for search/RAG? | +-- YES -> client.embeddings.create() with mistral-embed +-- Pre-configured agent? +-- YES -> client.agents.complete() with agent ID ``` ### Which Model to Choose ``` What is your task? +-- Most capable general purpose -> mistral-large-latest +-- Balanced cost/performance -> mistral-medium-latest +-- Fast + cost-efficient -> mistral-small-latest +-- Minimal / edge deployment -> ministral-3b-latest +-- Complex reasoning / math -> magistral-medium-latest +-- Code generation (chat) -> codestral-latest or devstral-latest +-- Code completion (FIM) -> codestral-latest +-- Vision / image analysis -> mistral-small-latest (or any vision-capable model) +-- Embeddings -> mistral-embed +-- Code embeddings -> codestral-embed-latest ``` ### Streaming vs Non-Streaming ``` Is the response user-facing? +-- YES -> Use client.chat.stream() | +-- Iterate with: for await (const event of result) | +-- Access content: event.data.choices[0]?.delta?.content +-- NO -> Use client.chat.complete() +-- Background processing -> chat.complete() +-- Structured output -> chat.parse() with Zod ``` ### When to Use This SDK vs a Provider-Agnostic SDK ``` Do you need multiple LLM providers (Mistral + others)? +-- YES -> Not this skill's scope -- use a unified provider SDK +-- NO -> Do you need Mistral-specific features? +-- YES -> Use Mistral SDK directly | Examples: Codestral FIM, Mistral Agents, | Voxtral audio, OCR, custom endpoints +-- NO -> Mistral SDK is simplest for Mistral-only use ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Using `response_format` (snake_case) instead of `responseFormat` (camelCase) -- silently ignored, no error thrown - Using `input` (singular) for embeddings instead of `inputs` (plural) -- Mistral-specific naming - Not configuring retries for production (SDK defaults to `strategy: "none"` -- zero retries) - Hardcoding API keys instead of using environment variables - Accessing `chunk.choices[0]?.delta?.content` directly on streaming events instead of `event.data.choices[0]?.delta?.content` **Medium Priority Issues:** - Not setting `timeoutMs` for production (default is `-1`, meaning no timeout -- requests can hang indefinitely) - Using `max_tokens` instead of `maxTokens` (camelCase SDK convention) - Missing `system` role message for behavior guidance - Using `tool_choice` instead of `toolChoice` - Using `tool_calls` instead of `toolCalls` when reading responses **Common Mistakes:** - Importing from `"mistralai"` instead of `"@mistralai/mistralai"` -- the correct package name has the org scope - Using CommonJS `require()` -- the package is ESM-only, use `import` or `await import()` - Confusing Mistral's `imageUrl: "url"` (flat string) with OpenAI's `image_url: { url: "..." }` (nested object) - Using `client.chat.completions.create()` (OpenAI pattern) instead of `client.chat.complete()` (Mistral pattern) - Assuming embedding dimensions match OpenAI's -- `mistral-embed` returns 1024-dimensional vectors, not 1536 **Gotchas & Edge Cases:** - The SDK is ESM-only. In CommonJS projects, you must use `const { Mistral } = await import("@mistralai/mistralai")`. - Streaming content may be `string | string[]` -- cast or check type when writing to stdout. - `chat.parse()` requires a Zod schema passed to `responseFormat` -- it does not accept `{ type: "json_object" }`. - The `apiKey` constructor option accepts a string OR an async function `() => Promise<string>` for dynamic key rotation. - Model aliases like `mistral-large-latest` resolve to the latest version of that model tier. Pin to specific versions (e.g., `mistral-large-3-25-12`) for reproducibility. - `toolChoice: "any"` forces the model to call a tool. `toolChoice: "auto"` lets the model decide. `toolChoice: "none"` prevents tool calls. - `parallelToolCalls: false` forces sequential tool calling (default `true` allows parallel). - FIM endpoint (`fim.complete()`) uses `prompt` + `suffix` parameters, NOT the `messages` array. - `safePrompt: true` injects Mistral's safety system prompt before your messages. - The SDK provides standalone functions (e.g., `chatComplete()` from `"@mistralai/mistralai/funcs/chatComplete.js"`) for tree-shaking in browser/edge runtimes. </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST use `responseFormat` (camelCase) in SDK calls -- NOT `response_format` (snake_case). The SDK uses camelCase property names throughout.)** **(You MUST configure retries explicitly -- the SDK defaults to `strategy: "none"` (no retries), unlike OpenAI's SDK which retries automatically)** **(You MUST consume streaming results with `for await (const event of result)` and access content via `event.data.choices[0]?.delta?.content` -- the event shape differs from OpenAI)** **(You MUST never hardcode API keys -- use `process.env["MISTRAL_API_KEY"]` with the bracket notation the SDK documents)** **(You MUST use `client.chat.parse()` with a Zod schema for structured outputs -- NOT manual `JSON.parse()` on completion content)** **Failure to follow these rules will produce broken API calls (snake_case properties silently ignored), unreliable production services (no retries), or incorrectly parsed streaming data.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.