ai-provider-openai-sdk
Official OpenAI SDK patterns for TypeScript/Node.js — client setup, Chat Completions, Responses API, streaming, structured outputs, function calling, embeddings, vision, audio, and production best practices
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-provider-openai-sdk/skills/ai-provider-openai-sdk
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
OpenAI SDK Patterns
Quick Guide: Use the official
openainpm package (v6+) to interact with OpenAI's API directly. Useclient.responses.create()(Responses API) for new projects with built-in tools and server-side state, orclient.chat.completions.create()(Chat Completions) for stateless chat flows. UsezodResponseFormatandclient.chat.completions.parse()for structured outputs. Use.stream()orstream: truefor streaming. Supports GPT-5.x family, GPT-4o, o4-mini, embeddings, vision, audio, and batch processing.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST use the Responses API (client.responses.create()) for new projects -- it provides better performance, built-in tools, and server-side conversation state)
(You MUST use zodResponseFormat() from openai/helpers/zod for structured outputs -- do NOT manually construct JSON schemas)
(You MUST handle errors using OpenAI.APIError and its subclasses -- never use bare catch blocks without error type checking)
(You MUST configure appropriate retries and timeouts for production use -- the SDK retries 2 times by default on 429/5xx errors)
(You MUST never hardcode API keys -- always use environment variables via process.env.OPENAI_API_KEY)
</critical_requirements>
Auto-detection: OpenAI, openai, client.chat.completions, client.responses.create, client.responses.parse, client.embeddings, client.audio, zodResponseFormat, zodTextFormat, zodFunction, zodResponsesFunction, runTools, GPT-5, GPT-4o, o4-mini, gpt-5-mini, text-embedding-3, whisper, tts, OPENAI_API_KEY, toFile
When to use:
- Building applications that call OpenAI models directly (GPT-5.x, GPT-4o, o4-mini, etc.)
- Implementing chat completions with streaming responses
- Using the Responses API for agentic workflows with built-in tools (web search, file search, code interpreter)
- Extracting structured data from LLM responses with Zod schema validation
- Implementing function calling / tool use with the Chat Completions or Responses API
- Creating embeddings for RAG pipelines or semantic search
- Processing images with vision models or audio with Whisper/TTS
- Running batch jobs for high-volume, cost-efficient processing
Key patterns covered:
- Client initialization and configuration (retries, timeouts, proxies)
- Chat Completions API (messages, streaming, function calling)
- Responses API (input, instructions, built-in tools, server-side state)
- Structured outputs with
zodResponseFormatandclient.chat.completions.parse() - Streaming with
for await...of,.stream()helper, and event handling - Embeddings API (
text-embedding-3-small,text-embedding-3-large) - Vision (image URLs, base64), Audio (Whisper transcription, TTS), Batch API
- Error handling, retries, timeouts, and production best practices
When NOT to use:
- Multi-provider applications where you need to switch between OpenAI, Anthropic, Google, etc. -- use a unified provider SDK instead
- React-specific chat UI hooks (
useChat,useCompletion) -- use a framework-integrated AI SDK - When you need a higher-level abstraction over multiple LLM providers
Examples Index
- Core: Setup & Configuration -- Client init, production config, Azure, error handling, request overrides
- Chat Completions -- Basic chat, multi-turn, token tracking, output length control
- Streaming --
stream: true,.stream()helper, Responses API streaming, abort - Tool/Function Calling -- Manual tools,
zodFunction,runToolsautomation, Responses API tools - Structured Output --
zodResponseFormat,zodTextFormat, refusal handling - Embeddings, Vision & Audio -- Semantic search, image analysis, transcription, TTS, batch processing
- Quick API Reference -- Model IDs, method signatures, error types, streaming events
<decision_framework>
Decision Framework
Which API to Use
Building a new application?
+-- YES -> Need built-in tools (web search, file search, code interpreter)?
| +-- YES -> Use Responses API (client.responses.create())
| +-- NO -> Need server-side conversation state?
| +-- YES -> Use Responses API with store: true
| +-- NO -> Either API works, prefer Responses for new code
+-- Existing Chat Completions code?
+-- Working fine? -> Keep using Chat Completions (fully supported)
+-- Need new features? -> Consider migrating to Responses API
Which Model to Choose
What is your task?
+-- General text generation -> gpt-5.4 (most capable) or gpt-4o (lower cost)
+-- Fast + cheap simple tasks -> gpt-5-mini or gpt-5-nano
+-- Complex reasoning / math -> gpt-5.4 or o4-mini
+-- Structured output -> gpt-5.4 or gpt-4o (best schema adherence)
+-- Vision (images) -> gpt-5.4 or gpt-4o
+-- Embeddings -> text-embedding-3-small (default) or text-embedding-3-large
+-- Transcription -> whisper-1 or gpt-4o-transcribe
+-- Text-to-speech -> tts-1 (fast) or gpt-4o-mini-tts (voice instructions)
+-- Batch processing -> gpt-5-mini (cheapest at 50% batch discount)
Streaming vs Non-Streaming
Is the response user-facing?
+-- YES -> Use streaming (stream: true or .stream())
| +-- Need event-level control? -> .stream() with event handlers
| +-- Simple text output? -> stream: true with for await
+-- NO -> Use non-streaming
+-- Background processing -> client.chat.completions.create()
+-- Structured output -> client.chat.completions.parse()
+-- High volume -> Batch API
When to Use This SDK vs a Provider-Agnostic SDK
Do you need multiple LLM providers (OpenAI + others)?
+-- YES -> Not this skill's scope -- use a unified provider SDK
+-- NO -> Do you need OpenAI-specific features?
+-- YES -> Use OpenAI SDK directly
| Examples: Responses API, Batch API,
| Realtime API, built-in web search/file search
+-- NO -> OpenAI SDK is simplest for OpenAI-only use
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Hardcoding API keys instead of using environment variables (security breach risk)
- Using bare
catchblocks without checkingOpenAI.APIError(hides API errors) - Not consuming streams returned by
stream: true(tokens are silently lost) - Using
JSON.parse()on completion content withoutzodResponseFormat(fragile, no validation) - Sending full conversation history every request when Responses API's
previous_response_idcould manage state
Medium Priority Issues:
- Not setting
maxRetries/timeoutfor production deployments (10 min default timeout may be too long) - Missing
developerrole message (no system instruction = unpredictable output style) - Using deprecated
systemrole instead ofdeveloperrole in Chat Completions - Not checking
finish_reasonfor'length'truncation - Ignoring
usagedata (no cost visibility)
Common Mistakes:
- Confusing Responses API (
client.responses.create()) with Chat Completions (client.chat.completions.create()) parameters -- they use different shapes - Using
messagesparameter with Responses API (it usesinputandinstructions) - Using
response_formatwith models that don't support structured outputs (need gpt-4o or later) - Using
max_tokenswith reasoning models (o4-mini, gpt-5.x) -- usemax_completion_tokensinstead - Not handling the case where
completion.choices[0].message.tool_callsis undefined - Forgetting that
runTools()defaults to max 10 completions -- setmaxChatCompletionsexplicitly
Gotchas & Edge Cases:
- The SDK auto-retries on 429 (rate limit) and 5xx errors -- 2 retries by default. Disable with
maxRetries: 0if you handle retries yourself. stream: truereturns raw SSE chunks. Use.stream()helper for a nicer event-based API.client.chat.completions.parse()throwsLengthFinishReasonErroriffinish_reasonis'length'andContentFilterFinishReasonErrorif'content_filter'.- Embedding responses return
Array<number>(the SDK requests base64 by default and decodes via Float32 internally for performance). No conversion needed -- you get a plain number array. - File uploads support
ReadStream,File,fetch()Response, andtoFile()helper -- use whichever matches your data source. - The Responses API's
store: trueenables server-side state but also means OpenAI stores your conversations. Setstore: falsefor sensitive data. developerrole replacessystemrole in newer models (gpt-4o and later).- Batch API has a 24h completion window and 50,000 request limit per batch.
- Audio transcription has a 25 MB file size limit.
- Zod schemas with
zodResponseFormatmust useadditionalProperties: false-- the SDK handles this automatically. zodTextFormatandzodResponseFormatare NOT compatible with Zod v4 -- use Zod v3.x until the SDK adds v4 support.- The Assistants API is deprecated (sunset August 2026) -- use the Responses API for new code.
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST use the Responses API (client.responses.create()) for new projects -- it provides better performance, built-in tools, and server-side conversation state)
(You MUST use zodResponseFormat() from openai/helpers/zod for structured outputs -- do NOT manually construct JSON schemas)
(You MUST handle errors using OpenAI.APIError and its subclasses -- never use bare catch blocks without error type checking)
(You MUST configure appropriate retries and timeouts for production use -- the SDK retries 2 times by default on 429/5xx errors)
(You MUST never hardcode API keys -- always use environment variables via process.env.OPENAI_API_KEY)
Failure to follow these rules will produce insecure, unreliable, or poorly-typed AI integrations.
</critical_reminders>
Files (skills)
-
examples
-
chat.md 3.1 KB
# OpenAI SDK -- Chat Completions Examples > Chat Completions API patterns: basic completion, multi-turn conversations, system/developer messages, temperature, token control. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio --- ## Basic Chat Completion ```typescript // basic-chat.ts import OpenAI from "openai"; const client = new OpenAI(); async function chat(userMessage: string): Promise<string> { const completion = await client.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: "You are a helpful assistant. Be concise.", }, { role: "user", content: userMessage }, ], }); const content = completion.choices[0].message.content; if (!content) { throw new Error("No content in response"); } console.log(`Tokens: ${completion.usage?.total_tokens}`); return content; } const answer = await chat("What is TypeScript in one sentence?"); console.log(answer); ``` --- ## Multi-Turn Conversations ```typescript import OpenAI from "openai"; import type { ChatCompletionMessageParam } from "openai/resources/chat/completions"; const client = new OpenAI(); const messages: ChatCompletionMessageParam[] = [ { role: "developer", content: "You are a TypeScript expert." }, { role: "user", content: "What is a union type?" }, ]; const completion = await client.chat.completions.create({ model: "gpt-4o", messages, }); // Append assistant response for next turn const assistantMessage = completion.choices[0].message; messages.push(assistantMessage); messages.push({ role: "user", content: "Give me a real-world example." }); const followUp = await client.chat.completions.create({ model: "gpt-4o", messages, }); ``` --- ## Token Usage Tracking ```typescript const completion = await client.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Hello" }], }); const usage = completion.usage; if (usage) { console.log(`Prompt tokens: ${usage.prompt_tokens}`); console.log(`Completion tokens: ${usage.completion_tokens}`); console.log(`Total tokens: ${usage.total_tokens}`); // Cached tokens (prompt caching) if (usage.prompt_tokens_details?.cached_tokens) { console.log(`Cached: ${usage.prompt_tokens_details.cached_tokens}`); } } ``` --- ## Controlling Output Length ```typescript const MAX_TOKENS = 500; const completion = await client.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Summarize this article." }], max_tokens: MAX_TOKENS, temperature: 0, // Deterministic output for caching }); if (completion.choices[0].finish_reason === "length") { console.warn("Output was truncated -- increase max_tokens"); } ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
core.md 4.5 KB
# OpenAI SDK -- Setup & Configuration Examples > Client initialization, environment config, production settings, Azure OpenAI, error handling, and abort patterns. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio --- ## Basic Client Setup ```typescript // lib/openai.ts import OpenAI from "openai"; // Reads OPENAI_API_KEY from env automatically const client = new OpenAI(); export { client }; ``` --- ## Production Configuration ```typescript // lib/openai.ts import OpenAI from "openai"; const TIMEOUT_MS = 30_000; const MAX_RETRIES = 3; const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY, timeout: TIMEOUT_MS, maxRetries: MAX_RETRIES, }); export { client }; ``` --- ## Azure OpenAI ```typescript // lib/azure-openai.ts import { AzureOpenAI } from "openai"; const azureClient = new AzureOpenAI({ apiKey: process.env.AZURE_OPENAI_API_KEY, apiVersion: "2024-10-21", endpoint: process.env.AZURE_OPENAI_ENDPOINT, }); export { azureClient }; ``` --- ## Production Error Handling ```typescript // error-handling.ts import OpenAI from "openai"; const TIMEOUT_MS = 30_000; const MAX_RETRIES = 3; const client = new OpenAI({ timeout: TIMEOUT_MS, maxRetries: MAX_RETRIES, }); async function safeCompletion(prompt: string): Promise<string | null> { try { const completion = await client.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: "You are a helpful assistant." }, { role: "user", content: prompt }, ], }); // Check for truncation if (completion.choices[0].finish_reason === "length") { console.warn("Response was truncated"); } return completion.choices[0].message.content; } catch (error) { if (error instanceof OpenAI.APIError) { console.error(`OpenAI API Error [${error.status}]: ${error.message}`); console.error(`Request ID: ${error.request_id}`); if (error instanceof OpenAI.RateLimitError) { console.error("Rate limited. SDK will auto-retry."); // If we get here, all retries were exhausted return null; } if (error instanceof OpenAI.AuthenticationError) { throw new Error( "Invalid API key. Check OPENAI_API_KEY environment variable.", ); } if (error instanceof OpenAI.BadRequestError) { console.error("Invalid request parameters:", error.message); return null; } // Server errors (5xx) -- SDK auto-retries, if we're here all retries failed if (error instanceof OpenAI.InternalServerError) { console.error("OpenAI server error after all retries"); return null; } } // Network/connection errors if (error instanceof OpenAI.APIConnectionError) { console.error("Network error:", error.message); return null; } // Unknown errors should be re-thrown throw error; } } const result = await safeCompletion("Hello!"); if (result) { console.log(result); } else { console.error("Failed to get completion"); } ``` --- ## Stream Error Handling ```typescript const stream = client.chat.completions.stream({ model: "gpt-4o", messages: [{ role: "user", content: "Hello" }], }); stream.on("error", (error) => { if (error instanceof OpenAI.APIError) { console.error(`Stream API error: ${error.status}`); } else { console.error("Stream connection error:", error); } }); stream.on("content", (delta) => { process.stdout.write(delta); }); await stream.finalContent(); ``` --- ## Request Cancellation with AbortController ```typescript const controller = new AbortController(); const ABORT_TIMEOUT_MS = 5_000; // Cancel after timeout setTimeout(() => controller.abort(), ABORT_TIMEOUT_MS); try { const completion = await client.chat.completions.create( { model: "gpt-4o", messages: [{ role: "user", content: "Hello" }] }, { signal: controller.signal }, ); } catch (error) { if (error instanceof Error && error.name === "AbortError") { console.log("Request was cancelled"); } } ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For per-request overrides, request ID tracking, and error type tables, see [reference.md](../reference.md)._ -
embeddings-vision-audio.md 9.7 KB
# OpenAI SDK -- Embeddings, Vision & Audio Examples > Embeddings for semantic search, vision/image inputs, audio transcription and TTS, file uploads, and batch processing. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod --- ## Embeddings and Semantic Search ```typescript // embeddings.ts import OpenAI from "openai"; const client = new OpenAI(); const EMBEDDING_MODEL = "text-embedding-3-small"; const SIMILARITY_THRESHOLD = 0.7; const TOP_K = 3; function cosineSimilarity(a: number[], b: number[]): number { let dot = 0; let normA = 0; let normB = 0; for (let i = 0; i < a.length; i++) { dot += a[i] * b[i]; normA += a[i] * a[i]; normB += b[i] * b[i]; } return dot / (Math.sqrt(normA) * Math.sqrt(normB)); } // Index documents const documents = [ "TypeScript provides static type checking for JavaScript.", "React is a library for building user interfaces.", "Node.js is a JavaScript runtime built on V8.", "PostgreSQL is a powerful relational database.", "Docker containers package applications with dependencies.", ]; const docEmbeddings = await client.embeddings.create({ model: EMBEDDING_MODEL, input: documents, }); const indexedDocs = documents.map((text, i) => ({ text, embedding: Array.from(docEmbeddings.data[i].embedding), })); // Search async function search( query: string, ): Promise<Array<{ text: string; score: number }>> { const queryEmbedding = await client.embeddings.create({ model: EMBEDDING_MODEL, input: query, }); const queryVector = Array.from(queryEmbedding.data[0].embedding); return indexedDocs .map((doc) => ({ text: doc.text, score: cosineSimilarity(queryVector, doc.embedding), })) .filter((r) => r.score > SIMILARITY_THRESHOLD) .sort((a, b) => b.score - a.score) .slice(0, TOP_K); } const results = await search("What is TypeScript?"); results.forEach((r) => { console.log(`[${r.score.toFixed(3)}] ${r.text}`); }); ``` --- ## Reduced Embedding Dimensions ```typescript // Use fewer dimensions for cost/speed tradeoff const response = await client.embeddings.create({ model: "text-embedding-3-large", input: "Some text to embed.", dimensions: 256, // Reduce from default 3072 }); ``` --- ## Vision -- Image from URL ```typescript // vision.ts import OpenAI from "openai"; const client = new OpenAI(); async function analyzeImageUrl( imageUrl: string, question: string, ): Promise<string> { const response = await client.chat.completions.create({ model: "gpt-4o", messages: [ { role: "user", content: [ { type: "text", text: question }, { type: "image_url", image_url: { url: imageUrl } }, ], }, ], }); return response.choices[0].message.content ?? ""; } ``` --- ## Vision -- Local Image (Base64) ```typescript import OpenAI from "openai"; import { readFileSync } from "node:fs"; const client = new OpenAI(); async function analyzeLocalImage( imagePath: string, question: string, ): Promise<string> { const imageBuffer = readFileSync(imagePath); const base64Image = imageBuffer.toString("base64"); const mimeType = imagePath.endsWith(".png") ? "image/png" : "image/jpeg"; const response = await client.chat.completions.create({ model: "gpt-4o", messages: [ { role: "user", content: [ { type: "text", text: question }, { type: "image_url", image_url: { url: `data:${mimeType};base64,${base64Image}`, detail: "high", // 'low' | 'high' | 'auto' }, }, ], }, ], }); return response.choices[0].message.content ?? ""; } ``` --- ## Vision -- Multiple Images ```typescript const response = await client.chat.completions.create({ model: "gpt-4o", messages: [ { role: "user", content: [ { type: "text", text: "Compare these two images." }, { type: "image_url", image_url: { url: "https://example.com/image1.jpg" }, }, { type: "image_url", image_url: { url: "https://example.com/image2.jpg" }, }, ], }, ], }); ``` --- ## Audio Transcription (Speech-to-Text) ```typescript // audio.ts import OpenAI from "openai"; import { createReadStream } from "node:fs"; const client = new OpenAI(); // Basic transcription async function transcribe(audioPath: string): Promise<string> { const transcription = await client.audio.transcriptions.create({ model: "whisper-1", file: createReadStream(audioPath), language: "en", }); return transcription.text; } // With word-level timestamps async function transcribeWithTimestamps(audioPath: string) { const transcription = await client.audio.transcriptions.create({ model: "whisper-1", file: createReadStream(audioPath), response_format: "verbose_json", timestamp_granularities: ["word"], }); return transcription.words?.map((w) => ({ word: w.word, start: w.start, end: w.end, })); } const transcript = await transcribe("./recording.mp3"); console.log("Transcript:", transcript); ``` --- ## Text-to-Speech (TTS) ```typescript import OpenAI from "openai"; import { writeFileSync } from "node:fs"; const client = new OpenAI(); // Basic TTS async function textToSpeech( text: string, outputPath: string, voice: | "alloy" | "ash" | "ballad" | "coral" | "echo" | "fable" | "nova" | "onyx" | "sage" | "shimmer" = "alloy", ): Promise<void> { const speech = await client.audio.speech.create({ model: "tts-1", voice, input: text, }); const buffer = Buffer.from(await speech.arrayBuffer()); writeFileSync(outputPath, buffer); console.log(`Speech saved to ${outputPath}`); } // With voice instructions (gpt-4o-mini-tts only) async function textToSpeechWithStyle( text: string, outputPath: string, style: string, ): Promise<void> { const speech = await client.audio.speech.create({ model: "gpt-4o-mini-tts", voice: "coral", input: text, instructions: style, }); const buffer = Buffer.from(await speech.arrayBuffer()); writeFileSync(outputPath, buffer); } await textToSpeech("Hello, welcome to the demo!", "./output.mp3", "nova"); await textToSpeechWithStyle( "Breaking news from the tech world!", "./news.mp3", "Speak like an excited news anchor", ); ``` --- ## File Uploads ```typescript import OpenAI, { toFile } from "openai"; import { createReadStream } from "node:fs"; const client = new OpenAI(); // From file path (Node.js ReadStream) const fileFromPath = await client.files.create({ file: createReadStream("training-data.jsonl"), purpose: "fine-tune", }); // From Buffer using toFile helper const fileFromBuffer = await client.files.create({ file: await toFile(Buffer.from('{"prompt": "Hi"}'), "data.jsonl"), purpose: "fine-tune", }); // From fetch Response const fileFromFetch = await client.files.create({ file: await fetch("https://example.com/data.jsonl"), purpose: "fine-tune", }); console.log(`Uploaded: ${fileFromPath.id}`); ``` --- ## Batch Processing ```typescript // batch-processing.ts import OpenAI, { toFile } from "openai"; const client = new OpenAI(); const POLL_INTERVAL_MS = 30_000; interface BatchRequest { custom_id: string; method: "POST"; url: string; body: { model: string; messages: Array<{ role: string; content: string }>; }; } // Create batch input function createBatchInput(prompts: string[]): string { const requests: BatchRequest[] = prompts.map((prompt, index) => ({ custom_id: `req-${index}`, method: "POST", url: "/v1/chat/completions", body: { model: "gpt-4o-mini", messages: [ { role: "developer", content: "Classify the sentiment as positive, negative, or neutral.", }, { role: "user", content: prompt }, ], }, })); return requests.map((r) => JSON.stringify(r)).join("\n"); } async function runBatch(prompts: string[]): Promise<string> { // Upload input file const jsonl = createBatchInput(prompts); const inputFile = await client.files.create({ file: await toFile(Buffer.from(jsonl), "batch-input.jsonl"), purpose: "batch", }); // Create batch const batch = await client.batches.create({ input_file_id: inputFile.id, endpoint: "/v1/chat/completions", completion_window: "24h", }); console.log(`Batch ${batch.id} created. Polling...`); // Poll for completion let status = batch; while ( !["completed", "failed", "cancelled", "expired"].includes(status.status) ) { await new Promise((resolve) => setTimeout(resolve, POLL_INTERVAL_MS)); status = await client.batches.retrieve(batch.id); console.log( `Status: ${status.status} (${status.request_counts?.completed ?? 0}/${status.request_counts?.total ?? 0})`, ); } if (status.status !== "completed" || !status.output_file_id) { throw new Error(`Batch ${status.status}: ${JSON.stringify(status.errors)}`); } // Download results const outputFile = await client.files.content(status.output_file_id); return outputFile.text(); } const prompts = [ "I love this product! It exceeded my expectations.", "The service was terrible and the food was cold.", "The meeting was rescheduled to next Tuesday.", ]; const results = await runBatch(prompts); console.log("Batch results:", results); ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
openai-sdk.md 552 B
# Deprecated -- Examples Have Been Split This file has been replaced by topic-specific example files: - [core.md](core.md) -- Client setup, error handling, configuration - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio, files, batch See [SKILL.md](../SKILL.md) for the examples index. -
streaming.md 2.9 KB
# OpenAI SDK -- Streaming Examples > Streaming patterns for Chat Completions and Responses API: `stream: true` with async iterators, `.stream()` event-based helper, Responses API streaming. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat Completions API - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio --- ## Basic Streaming with `for await` ```typescript // streaming-chat.ts import OpenAI from "openai"; const client = new OpenAI(); // Chat Completions streaming const stream = await client.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: "You are a helpful assistant." }, { role: "user", content: "Explain async/await in TypeScript." }, ], stream: true, }); for await (const chunk of stream) { const content = chunk.choices[0]?.delta?.content; if (content) { process.stdout.write(content); } } console.log(); // newline ``` --- ## Event-Based Streaming with `.stream()` Helper ```typescript import OpenAI from "openai"; const client = new OpenAI(); async function streamWithEvents(prompt: string): Promise<string> { const stream = client.chat.completions.stream({ model: "gpt-4o", messages: [ { role: "developer", content: "You are a helpful assistant." }, { role: "user", content: prompt }, ], }); stream.on("content", (delta) => { process.stdout.write(delta); }); stream.on("error", (error) => { console.error("Stream error:", error); }); const content = await stream.finalContent(); console.log(); // newline return content ?? ""; } const result = await streamWithEvents("Explain promises in JavaScript."); console.log("Final result length:", result.length); ``` --- ## Responses API Streaming ```typescript import OpenAI from "openai"; const client = new OpenAI(); const stream = await client.responses.create({ model: "gpt-4o", input: "Write a haiku about TypeScript.", stream: true, }); for await (const event of stream) { if (event.type === "response.output_text.delta") { process.stdout.write(event.delta); } } console.log(); ``` --- ## Stream Abort ```typescript const stream = client.chat.completions.stream({ model: "gpt-4o", messages: [{ role: "user", content: "Tell me a long story." }], }); // Cancel after a timeout const ABORT_TIMEOUT_MS = 5_000; setTimeout(() => stream.abort(), ABORT_TIMEOUT_MS); stream.on("content", (delta) => { process.stdout.write(delta); }); try { await stream.finalContent(); } catch (error) { console.log("Stream aborted"); } ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For stream method signatures, event types, and API tables, see [reference.md](../reference.md)._ -
structured-output.md 4.4 KB
# OpenAI SDK -- Structured Output Examples > Type-safe structured responses with Zod: `zodResponseFormat` for Chat Completions, `zodTextFormat` for Responses API, refusal handling, complex schemas. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio --- ## Chat Completions Structured Output with `zodResponseFormat` ```typescript // structured-output.ts import OpenAI from "openai"; import { zodResponseFormat } from "openai/helpers/zod"; import { z } from "zod"; const client = new OpenAI(); // Define the schema for extracted data const ArticleSummary = z.object({ title: z.string(), summary: z.string(), keyPoints: z.array(z.string()), sentiment: z.enum(["positive", "negative", "neutral"]), wordCount: z.number(), }); type ArticleSummary = z.infer<typeof ArticleSummary>; async function extractArticleSummary( articleText: string, ): Promise<ArticleSummary | null> { const completion = await client.chat.completions.parse({ model: "gpt-4o", messages: [ { role: "developer", content: "Extract a structured summary from the provided article text.", }, { role: "user", content: articleText }, ], response_format: zodResponseFormat(ArticleSummary, "article_summary"), }); const message = completion.choices[0].message; // Handle safety refusals if (message.refusal) { console.warn("Model refused:", message.refusal); return null; } return message.parsed; } const article = ` TypeScript 5.8 brings exciting new features including improved type inference, better error messages, and performance optimizations. The release focuses on developer experience improvements that make everyday coding more productive. `; const summary = await extractArticleSummary(article); if (summary) { console.log(`Title: ${summary.title}`); console.log(`Sentiment: ${summary.sentiment}`); console.log("Key Points:"); summary.keyPoints.forEach((point) => console.log(` - ${point}`)); } ``` --- ## Handling Refusals ```typescript import OpenAI from "openai"; import { zodResponseFormat } from "openai/helpers/zod"; import { z } from "zod"; const client = new OpenAI(); const SomeSchema = z.object({ content: z.string(), }); const completion = await client.chat.completions.parse({ model: "gpt-4o", messages: [{ role: "user", content: "Generate harmful content" }], response_format: zodResponseFormat(SomeSchema, "output"), }); const message = completion.choices[0].message; if (message.refusal) { console.log("Model refused:", message.refusal); } else if (message.parsed) { console.log("Parsed output:", message.parsed); } ``` --- ## Responses API Structured Output with `zodTextFormat` ```typescript import OpenAI from "openai"; import { zodTextFormat } from "openai/helpers/zod"; import { z } from "zod"; const client = new OpenAI(); const PersonSchema = z.object({ name: z.string(), age: z.number(), }); const response = await client.responses.parse({ model: "gpt-4o", input: "Jane is 54 years old.", text: { format: zodTextFormat(PersonSchema, "person"), }, }); console.log(response.output_parsed); // { name: 'Jane', age: 54 } ``` --- ## Manual JSON Schema (Anti-Pattern) Avoid this -- use `zodResponseFormat` instead: ```typescript // BAD: Manually constructing JSON schema const completion = await client.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Extract data" }], response_format: { type: "json_schema", json_schema: { name: "event", strict: true, schema: { type: "object", properties: { name: { type: "string" }, date: { type: "string" }, }, required: ["name", "date"], additionalProperties: false, }, }, }, }); // Then manually JSON.parse the content -- error-prone const data = JSON.parse(completion.choices[0].message.content ?? "{}"); ``` **Why bad:** Manual JSON schema is verbose and error-prone, no type safety, manual parsing can fail. Use `zodResponseFormat` + `.parse()` instead. --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
tools.md 7.5 KB
# OpenAI SDK -- Tool/Function Calling Examples > Function calling patterns: manual tool definitions, Zod-based tools with `zodFunction`, automated tool loops with `runTools`, Responses API function calling. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio --- ## Chat Completions with Manual Tool Definitions ```typescript import OpenAI from "openai"; import type { ChatCompletionTool } from "openai/resources/chat/completions"; const client = new OpenAI(); const tools: ChatCompletionTool[] = [ { type: "function", function: { name: "get_weather", description: "Get the current weather for a location", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, unit: { type: "string", enum: ["celsius", "fahrenheit"] }, }, required: ["location"], additionalProperties: false, }, strict: true, }, }, ]; const completion = await client.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "What is the weather in Tokyo?" }], tools, }); const toolCall = completion.choices[0].message.tool_calls?.[0]; if (toolCall) { const args = JSON.parse(toolCall.function.arguments); console.log(`Call ${toolCall.function.name} with:`, args); } ``` --- ## Zod-Based Tools with `zodFunction` ```typescript // function-calling.ts import OpenAI from "openai"; import { zodFunction } from "openai/helpers/zod"; import { z } from "zod"; const client = new OpenAI(); // Define tool schemas with Zod const GetWeatherParams = z.object({ location: z.string().describe('City name, e.g. "San Francisco"'), unit: z.enum(["celsius", "fahrenheit"]).default("celsius"), }); const SearchDatabaseParams = z.object({ query: z.string().describe("Search query string"), limit: z.number().default(10).describe("Max results to return"), }); // Tool implementations async function getWeather( args: z.infer<typeof GetWeatherParams>, ): Promise<string> { // In production, call a real weather API return JSON.stringify({ location: args.location, temperature: 22, unit: args.unit, condition: "sunny", }); } async function searchDatabase( args: z.infer<typeof SearchDatabaseParams>, ): Promise<string> { // In production, query your database return JSON.stringify({ results: [{ id: 1, title: `Result for: ${args.query}` }], total: 1, }); } // Map tool names to implementations const toolImplementations: Record<string, (args: unknown) => Promise<string>> = { get_weather: getWeather as (args: unknown) => Promise<string>, search_database: searchDatabase as (args: unknown) => Promise<string>, }; // Create completion with tools const completion = await client.chat.completions.parse({ model: "gpt-4o", messages: [ { role: "developer", content: "You help users by calling available tools.", }, { role: "user", content: "What is the weather in Tokyo?" }, ], tools: [ zodFunction({ name: "get_weather", parameters: GetWeatherParams }), zodFunction({ name: "search_database", parameters: SearchDatabaseParams }), ], }); // Process tool calls const message = completion.choices[0].message; if (message.tool_calls && message.tool_calls.length > 0) { for (const toolCall of message.tool_calls) { const fnName = toolCall.function.name; const args = JSON.parse(toolCall.function.arguments); console.log(`Calling ${fnName} with:`, args); const impl = toolImplementations[fnName]; if (impl) { const result = await impl(args); console.log(`Result: ${result}`); } } } ``` --- ## Automated Tool Loop with `runTools` ```typescript // run-tools.ts import OpenAI from "openai"; const client = new OpenAI(); const MAX_TOOL_CALLS = 5; async function getWeather(args: { location: string }): Promise<string> { return `Weather in ${args.location}: 22C, sunny`; } async function getTime(args: { timezone: string }): Promise<string> { return `Time in ${args.timezone}: ${new Date().toLocaleTimeString()}`; } const runner = client.chat.completions.runTools({ model: "gpt-4o", messages: [ { role: "developer", content: "Help users with weather and time queries." }, { role: "user", content: "What is the weather and current time in London?", }, ], tools: [ { type: "function", function: { function: getWeather, parse: JSON.parse, description: "Get current weather for a location", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, }, required: ["location"], }, }, }, { type: "function", function: { function: getTime, parse: JSON.parse, description: "Get current time in a timezone", parameters: { type: "object", properties: { timezone: { type: "string", description: "Timezone, e.g. Europe/London", }, }, required: ["timezone"], }, }, }, ], maxChatCompletions: MAX_TOOL_CALLS, }); // Monitor the tool execution loop runner.on("message", (msg) => { if (msg.role === "assistant" && msg.tool_calls) { console.log(`[Tool calls: ${msg.tool_calls.length}]`); } if (msg.role === "tool") { console.log(`[Tool result received]`); } }); const finalContent = await runner.finalContent(); console.log("\nFinal answer:", finalContent); ``` --- ## Responses API Function Calling ```typescript // responses-function-calling.ts import OpenAI from "openai"; const client = new OpenAI(); // Define the function tool for Responses API const response = await client.responses.create({ model: "gpt-4o", instructions: "You are a helpful assistant with access to weather data.", input: "What is the weather like in San Francisco and New York?", tools: [ { type: "function", name: "get_weather", description: "Get current weather for a city", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, }, required: ["location"], additionalProperties: false, }, }, ], }); // Process function call outputs const functionCalls = response.output.filter( (item) => item.type === "function_call", ); for (const call of functionCalls) { console.log(`Function: ${call.name}`); console.log(`Arguments: ${call.arguments}`); console.log(`Call ID: ${call.call_id}`); } // Submit function results back to continue the conversation if (functionCalls.length > 0) { const toolOutputs = functionCalls.map((call) => ({ type: "function_call_output" as const, call_id: call.call_id, output: JSON.stringify({ temperature: 22, condition: "sunny" }), })); const followUp = await client.responses.create({ model: "gpt-4o", input: toolOutputs, previous_response_id: response.id, store: true, }); console.log("Final answer:", followUp.output_text); } ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
-
-
reference.md 12.9 KB
# OpenAI SDK Quick Reference > Client configuration, model IDs, API methods, error types, and helper functions. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Package Installation ```bash # Core package (always required) npm install openai # For structured outputs (optional but recommended) npm install zod ``` --- ## Client Configuration ```typescript import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY, // Auto-reads from env if not set timeout: 30_000, // Request timeout in ms (default: 600_000 = 10 min) maxRetries: 3, // Retry count on 429/5xx (default: 2) baseURL: "https://api.openai.com/v1", // Override for proxies dangerouslyAllowBrowser: false, // Must be true for browser usage }); ``` ### Environment Variables | Variable | Purpose | | ----------------- | -------------------------- | | `OPENAI_API_KEY` | API key (auto-detected) | | `OPENAI_ORG_ID` | Organization ID (optional) | | `OPENAI_BASE_URL` | Custom base URL (optional) | --- ## Model IDs ### Language Models (Chat / Text Generation) | Model | Use Case | Context Window | Notes | | ------------- | ----------------------------------------- | -------------- | -------------------------------------- | | `gpt-5.4` | Most capable, general purpose + reasoning | 1M | Recommended for new projects | | `gpt-5-mini` | Cost-optimized, balanced speed/quality | 1M | Replaces gpt-4o-mini for most uses | | `gpt-5-nano` | High-throughput, straightforward tasks | 1M | Cheapest GPT-5 variant | | `gpt-4o` | General purpose, proven | 128K | Still available, lower cost than GPT-5 | | `gpt-4o-mini` | Fast, cheap, simple tasks | 128K | Still available, migrate to gpt-5-mini | | `o4-mini` | Fast reasoning | 200K | Use `max_completion_tokens` | ### Embedding Models | Model | Dimensions (default) | Max Input | | ------------------------ | -------------------- | ----------- | | `text-embedding-3-small` | 1536 | 8191 tokens | | `text-embedding-3-large` | 3072 | 8191 tokens | ### Audio Models | Model | Type | Use Case | | ------------------------ | ------------- | --------------------------- | | `whisper-1` | Transcription | General speech-to-text | | `gpt-4o-mini-transcribe` | Transcription | Fewer hallucinations | | `gpt-4o-transcribe` | Transcription | Higher accuracy | | `tts-1` | TTS | Fast text-to-speech | | `tts-1-hd` | TTS | High-quality text-to-speech | | `gpt-4o-mini-tts` | TTS | Voice instruction support | ### TTS Voices `alloy`, `ash`, `ballad`, `coral`, `echo`, `fable`, `nova`, `onyx`, `sage`, `shimmer` --- ## API Methods Reference ### Chat Completions API ```typescript // Standard completion const completion = await client.chat.completions.create({ model: "gpt-4o", // Required messages: [], // Required: ChatCompletionMessageParam[] temperature: 0.7, // 0-2 (default: 1) max_tokens: 1000, // Max output tokens (use max_completion_tokens for reasoning models) top_p: 1, // Nucleus sampling frequency_penalty: 0, // -2 to 2 presence_penalty: 0, // -2 to 2 tools: [], // ChatCompletionTool[] tool_choice: "auto", // 'auto' | 'required' | 'none' | { type: 'function', function: { name: string } } response_format: undefined, // zodResponseFormat() or { type: 'json_object' } stream: false, // Enable streaming stop: undefined, // string | string[] -- stop sequences seed: undefined, // number -- for reproducibility logprobs: false, // Include log probabilities user: undefined, // End-user ID for abuse monitoring }); // Structured output parsing const parsed = await client.chat.completions.parse({ model: "gpt-4o", messages: [], response_format: zodResponseFormat(schema, "name"), }); // Event-based streaming const stream = client.chat.completions.stream({ model: "gpt-4o", messages: [], }); // Automated tool execution const runner = client.chat.completions.runTools({ model: "gpt-4o", messages: [], tools: [], maxChatCompletions: 10, // Max tool call loops (default: 10) }); ``` ### Responses API ```typescript const response = await client.responses.create({ model: "gpt-4o", // Required input: "", // Required: string or InputItem[] instructions: "", // System instruction (replaces 'developer' role) tools: [], // Built-in or function tools previous_response_id: "", // Chain conversations store: true, // Enable server-side state stream: false, // Enable streaming text: { // Structured output config format: zodTextFormat(schema, "name"), }, }); // Access helpers response.output_text; // Direct text access response.output; // Array of typed output Items response.id; // Response ID for chaining ``` ### Embeddings ```typescript const response = await client.embeddings.create({ model: "text-embedding-3-small", // Required input: "", // string or string[] dimensions: undefined, // Reduce dimensions (optional) encoding_format: "float", // 'float' | 'base64' }); ``` ### Audio ```typescript // Transcription const transcription = await client.audio.transcriptions.create({ model: "whisper-1", // Required file: readStream, // Required: ReadStream or File language: "en", // ISO 639-1 code response_format: "json", // 'json' | 'text' | 'srt' | 'vtt' | 'verbose_json' temperature: 0, // 0-1 timestamp_granularities: [], // ['word', 'segment'] (verbose_json only) }); // Text-to-Speech const speech = await client.audio.speech.create({ model: "tts-1", // Required voice: "alloy", // Required input: "", // Required: text to speak response_format: "mp3", // mp3, opus, aac, flac, wav, pcm speed: 1.0, // 0.25-4.0 instructions: "", // Voice style (gpt-4o-mini-tts only) }); // Translation (to English) const translation = await client.audio.translations.create({ model: "whisper-1", file: readStream, }); ``` ### Files ```typescript // Upload const file = await client.files.create({ file: readStream, // ReadStream | File | Response | toFile() purpose: "fine-tune", // 'fine-tune' | 'batch' }); // List const files = await client.files.list(); // Retrieve const fileInfo = await client.files.retrieve("file-abc123"); // Download content const content = await client.files.content("file-abc123"); // Delete await client.files.del("file-abc123"); ``` ### Batch API ```typescript // Create batch const batch = await client.batches.create({ input_file_id: "file-abc123", endpoint: "/v1/chat/completions", // or '/v1/embeddings' completion_window: "24h", metadata: {}, }); // Retrieve status const status = await client.batches.retrieve("batch-abc123"); // List batches const batches = await client.batches.list(); // Cancel batch await client.batches.cancel("batch-abc123"); ``` --- ## Helper Functions ```typescript // Structured output with Zod import { zodResponseFormat } from "openai/helpers/zod"; import { zodTextFormat } from "openai/helpers/zod"; import { zodFunction } from "openai/helpers/zod"; import { zodResponsesFunction } from "openai/helpers/zod"; // Chat Completions structured output zodResponseFormat(zodSchema, "schema_name"); // Responses API structured output zodTextFormat(zodSchema, "schema_name"); // Chat Completions function tool from Zod schema zodFunction({ name: "tool_name", parameters: zodSchema }); // Responses API function tool from Zod schema zodResponsesFunction({ name: "tool_name", parameters: zodSchema }); // File conversion helper import { toFile } from "openai"; await toFile(buffer, "filename.ext"); ``` --- ## Error Types | Error Class | HTTP Status | Auto-Retried? | | --------------------------- | ----------- | ------------- | | `BadRequestError` | 400 | No | | `AuthenticationError` | 401 | No | | `PermissionDeniedError` | 403 | No | | `NotFoundError` | 404 | No | | `UnprocessableEntityError` | 422 | No | | `RateLimitError` | 429 | Yes | | `InternalServerError` | >= 500 | Yes | | `APIConnectionError` | N/A | Yes | | `APIConnectionTimeoutError` | N/A | Yes | All errors extend `OpenAI.APIError` with properties: - `.status` -- HTTP status code - `.message` -- Error message - `.request_id` -- Request ID for debugging - `.headers` -- Response headers --- ## Streaming Events (.stream() Helper) | Event | Arguments | Description | | ------------------------------------- | ----------------------------------------------------------------- | -------------------------------- | | `connect` | `()` | Connection established | | `chunk` | `(chunk, snapshot)` | Raw API chunk received | | `content` | `(delta, snapshot)` | Text content delta | | `content.delta` | `({ delta, snapshot, parsed })` | Content chunk with full snapshot | | `content.done` | `({ content, parsed })` | Content generation complete | | `refusal.delta` | `({ delta, snapshot })` | Refusal content delta | | `refusal.done` | `({ refusal })` | Refusal content complete | | `tool_calls.function.arguments.delta` | `({ name, index, arguments, arguments_delta, parsed_arguments })` | Tool argument streaming | | `tool_calls.function.arguments.done` | `({ name, index, arguments, parsed_arguments })` | Tool arguments complete | | `error` | `(error)` | Stream error | | `end` | `()` | Stream finished | ### Stream Methods ```typescript await stream.finalContent(); // Promise<string> -- last assistant content await stream.finalMessage(); // Promise<ChatCompletionMessage> await stream.allChatCompletions(); // Promise<ChatCompletion[]> stream.abort(); // Cancel stream and network request stream.controller; // Underlying AbortController ``` --- ## Responses API Event Types (Streaming) | Event Type | Description | | ---------------------------------------- | --------------------------- | | `response.created` | Response object created | | `response.output_text.delta` | Text content delta | | `response.output_text.done` | Text generation complete | | `response.function_call_arguments.delta` | Function argument streaming | | `response.function_call_arguments.done` | Function arguments complete | | `response.completed` | Full response complete | | `error` | Error occurred | --- ## Message Roles | Role | API | Description | | ----------- | ---------------- | ------------------------------------------------------ | | `developer` | Chat Completions | System instruction (replaces `system` in newer models) | | `system` | Chat Completions | Legacy system instruction | | `user` | Both | User input | | `assistant` | Chat Completions | Model response | | `tool` | Chat Completions | Tool result | --- ## Per-Request Overrides ```typescript // Override retries and timeout for a single request await client.chat.completions.create( { model: 'gpt-4o', messages: [...] }, { maxRetries: 5, timeout: 60_000, signal: abortController.signal, headers: { 'X-Custom-Header': 'value' }, }, ); ``` --- ## Request ID Tracking ```typescript // From response object completion._request_id; // From raw response headers const { data, response } = await client.chat.completions .create({ model: 'gpt-4o', messages: [...] }) .withResponse(); response.headers.get('x-request-id'); ``` -
SKILL.md 20.6 KB
--- name: ai-provider-openai-sdk description: Official OpenAI SDK patterns for TypeScript/Node.js — client setup, Chat Completions, Responses API, streaming, structured outputs, function calling, embeddings, vision, audio, and production best practices --- # OpenAI SDK Patterns > **Quick Guide:** Use the official `openai` npm package (v6+) to interact with OpenAI's API directly. Use `client.responses.create()` (Responses API) for new projects with built-in tools and server-side state, or `client.chat.completions.create()` (Chat Completions) for stateless chat flows. Use `zodResponseFormat` and `client.chat.completions.parse()` for structured outputs. Use `.stream()` or `stream: true` for streaming. Supports GPT-5.x family, GPT-4o, o4-mini, embeddings, vision, audio, and batch processing. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST use the Responses API (`client.responses.create()`) for new projects -- it provides better performance, built-in tools, and server-side conversation state)** **(You MUST use `zodResponseFormat()` from `openai/helpers/zod` for structured outputs -- do NOT manually construct JSON schemas)** **(You MUST handle errors using `OpenAI.APIError` and its subclasses -- never use bare catch blocks without error type checking)** **(You MUST configure appropriate retries and timeouts for production use -- the SDK retries 2 times by default on 429/5xx errors)** **(You MUST never hardcode API keys -- always use environment variables via `process.env.OPENAI_API_KEY`)** </critical_requirements> --- **Auto-detection:** OpenAI, openai, client.chat.completions, client.responses.create, client.responses.parse, client.embeddings, client.audio, zodResponseFormat, zodTextFormat, zodFunction, zodResponsesFunction, runTools, GPT-5, GPT-4o, o4-mini, gpt-5-mini, text-embedding-3, whisper, tts, OPENAI_API_KEY, toFile **When to use:** - Building applications that call OpenAI models directly (GPT-5.x, GPT-4o, o4-mini, etc.) - Implementing chat completions with streaming responses - Using the Responses API for agentic workflows with built-in tools (web search, file search, code interpreter) - Extracting structured data from LLM responses with Zod schema validation - Implementing function calling / tool use with the Chat Completions or Responses API - Creating embeddings for RAG pipelines or semantic search - Processing images with vision models or audio with Whisper/TTS - Running batch jobs for high-volume, cost-efficient processing **Key patterns covered:** - Client initialization and configuration (retries, timeouts, proxies) - Chat Completions API (messages, streaming, function calling) - Responses API (input, instructions, built-in tools, server-side state) - Structured outputs with `zodResponseFormat` and `client.chat.completions.parse()` - Streaming with `for await...of`, `.stream()` helper, and event handling - Embeddings API (`text-embedding-3-small`, `text-embedding-3-large`) - Vision (image URLs, base64), Audio (Whisper transcription, TTS), Batch API - Error handling, retries, timeouts, and production best practices **When NOT to use:** - Multi-provider applications where you need to switch between OpenAI, Anthropic, Google, etc. -- use a unified provider SDK instead - React-specific chat UI hooks (`useChat`, `useCompletion`) -- use a framework-integrated AI SDK - When you need a higher-level abstraction over multiple LLM providers --- ## Examples Index - [Core: Setup & Configuration](examples/core.md) -- Client init, production config, Azure, error handling, request overrides - [Chat Completions](examples/chat.md) -- Basic chat, multi-turn, token tracking, output length control - [Streaming](examples/streaming.md) -- `stream: true`, `.stream()` helper, Responses API streaming, abort - [Tool/Function Calling](examples/tools.md) -- Manual tools, `zodFunction`, `runTools` automation, Responses API tools - [Structured Output](examples/structured-output.md) -- `zodResponseFormat`, `zodTextFormat`, refusal handling - [Embeddings, Vision & Audio](examples/embeddings-vision-audio.md) -- Semantic search, image analysis, transcription, TTS, batch processing - [Quick API Reference](reference.md) -- Model IDs, method signatures, error types, streaming events --- <philosophy> ## Philosophy The official OpenAI SDK provides **direct, low-level access** to OpenAI's full API surface. It is the thinnest possible wrapper over the REST API, auto-generated from OpenAI's OpenAPI specification using Stainless. **Core principles:** 1. **Direct API access** -- No abstractions or provider layers. You get the exact API that OpenAI documents, with full TypeScript types. Every API feature is available immediately when OpenAI releases it. 2. **Two API paradigms** -- The **Responses API** (`client.responses.create()`) is the newer, recommended API with built-in tools and server-side state. The **Chat Completions API** (`client.chat.completions.create()`) remains fully supported for stateless chat flows. 3. **Built-in resilience** -- The SDK handles retries (2 by default on 429/5xx), timeouts (10 min default), and auto-pagination out of the box. 4. **Streaming as a first-class pattern** -- Use `stream: true` for SSE-based streaming, `.stream()` helper for event-based consumption, or `for await...of` for simple iteration. 5. **Type-safe structured outputs** -- `zodResponseFormat()` and `client.chat.completions.parse()` convert Zod schemas to JSON Schema and parse responses, giving you validated, typed objects. **When to use the OpenAI SDK directly:** - You only use OpenAI models and want the simplest, most direct integration - You need access to OpenAI-specific features (Responses API, Batch, Realtime) - You want minimal dependencies and zero abstraction overhead - You need the latest API features on day one **When NOT to use:** - You need to switch between providers (OpenAI, Anthropic, Google) -- use a unified provider SDK - You want React-specific chat UI hooks -- use a framework-integrated AI SDK - You want a higher-level agent framework -- consider OpenAI Agents SDK (`@openai/agents`) </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Setup Initialize the OpenAI client. It auto-reads `OPENAI_API_KEY` from the environment. ```typescript // lib/openai.ts -- basic setup import OpenAI from "openai"; const client = new OpenAI(); export { client }; ``` ```typescript // lib/openai.ts -- production configuration const TIMEOUT_MS = 30_000; const MAX_RETRIES = 3; const client = new OpenAI({ timeout: TIMEOUT_MS, maxRetries: MAX_RETRIES }); ``` **Why good:** Minimal setup, env var auto-detected, named constants for production settings **See:** [examples/core.md](examples/core.md) for Azure OpenAI, per-request overrides, error handling patterns --- ### Pattern 2: Chat Completions API Stateless text generation. You manage conversation history. ```typescript const completion = await client.chat.completions.create({ model: "gpt-4o", messages: [ { role: "developer", content: "You are a helpful coding assistant." }, { role: "user", content: "Explain TypeScript generics." }, ], }); console.log(completion.choices[0].message.content); ``` **Why good:** Clear message roles, `developer` message for system instructions, direct content access ```typescript // BAD: No developer message, no error handling const res = await client.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "do something" }], }); ``` **Why bad:** No system instruction means unpredictable behavior, vague prompt **See:** [examples/chat.md](examples/chat.md) for multi-turn, token tracking, output length control --- ### Pattern 3: Responses API (Recommended for New Projects) Newer API with built-in tools, server-side state, and better performance with reasoning models. ```typescript const response = await client.responses.create({ model: "gpt-4o", instructions: "You are a coding assistant.", input: "What are TypeScript generics?", }); console.log(response.output_text); ``` **Why good:** Clean separation of instructions and input, `output_text` helper, simpler than messages array ```typescript // BAD: Using Chat Completions parameters with Responses API const response = await client.responses.create({ model: "gpt-4o", messages: [{ role: "user", content: "Hello" }], // WRONG: use 'input' }); ``` **Why bad:** Responses API uses `input` and `instructions`, not `messages` #### Built-in Tools Web search (`{ type: "web_search_preview" }`), file search (`{ type: "file_search" }`), code interpreter (`{ type: "code_interpreter" }`). Chain conversations with `previous_response_id` and `store: true`. **See:** [examples/tools.md](examples/tools.md) for Responses API function calling with tool outputs --- ### Pattern 4: Streaming Use streaming for user-facing responses. ```typescript // Chat Completions -- stream: true with for-await const stream = await client.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Explain async/await." }], stream: true, }); for await (const chunk of stream) { const content = chunk.choices[0]?.delta?.content; if (content) process.stdout.write(content); } ``` ```typescript // Event-based with .stream() helper const stream = client.chat.completions.stream({ model: "gpt-4o", messages: [{ role: "user", content: "Tell me a story." }], }); stream.on("content", (delta) => process.stdout.write(delta)); const finalContent = await stream.finalContent(); ``` **Why good:** Progressive output for better UX, event-based API for granular control ```typescript // BAD: Not consuming the stream const stream = await client.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Hello" }], stream: true, }); // Stream never consumed -- tokens are lost ``` **Why bad:** Stream must be consumed via iteration or event handlers, otherwise tokens are lost **See:** [examples/streaming.md](examples/streaming.md) for Responses API streaming, abort, stream methods --- ### Pattern 5: Structured Outputs with Zod Use `zodResponseFormat()` and `.parse()` for type-safe structured responses. ```typescript import { zodResponseFormat } from "openai/helpers/zod"; import { z } from "zod"; const CalendarEvent = z.object({ name: z.string(), date: z.string(), participants: z.array(z.string()), }); const completion = await client.chat.completions.parse({ model: "gpt-4o", messages: [ { role: "developer", content: "Extract event details." }, { role: "user", content: "Alice and Bob meet next Tuesday for lunch." }, ], response_format: zodResponseFormat(CalendarEvent, "calendar_event"), }); const event = completion.choices[0].message.parsed; // Fully typed ``` **Why good:** Auto-converts schema, validates output, fully typed result, handles refusals **See:** [examples/structured-output.md](examples/structured-output.md) for Responses API (`zodTextFormat`), refusal handling, complex schemas --- ### Pattern 6: Function Calling / Tool Use Define functions the model can call. Use `zodFunction()` for type-safe definitions. ```typescript import { zodFunction } from "openai/helpers/zod"; import { z } from "zod"; const GetWeatherParams = z.object({ location: z.string().describe("City name"), unit: z.enum(["celsius", "fahrenheit"]).default("celsius"), }); const completion = await client.chat.completions.parse({ model: "gpt-4o", messages: [{ role: "user", content: "Weather in Paris?" }], tools: [zodFunction({ name: "get_weather", parameters: GetWeatherParams })], }); const toolCall = completion.choices[0].message.tool_calls?.[0]; if (toolCall?.type === "function") { console.log(toolCall.function.parsed_arguments); // Typed from Zod } ``` **Why good:** `zodFunction` provides type-safe argument parsing, `.describe()` guides the model Use `runTools()` for automated tool execution loops that handle the call-respond cycle automatically. **See:** [examples/tools.md](examples/tools.md) for `runTools`, manual tool definitions, Responses API function calling --- ### Pattern 7: Embeddings, Vision & Audio - **Embeddings:** `client.embeddings.create({ model: "text-embedding-3-small", input: [...] })` -- batch multiple inputs in one call - **Vision:** Multi-part content array with `{ type: "image_url", image_url: { url } }` for URL or base64 images - **Audio:** `client.audio.transcriptions.create()` for speech-to-text, `client.audio.speech.create()` for TTS - **Files:** `client.files.create()` with `ReadStream`, `Buffer` (via `toFile`), or `fetch()` Response - **Batch API:** Upload JSONL, create batch with `client.batches.create()`, poll for completion at 50% cost **See:** [examples/embeddings-vision-audio.md](examples/embeddings-vision-audio.md) for full examples with cosine similarity, base64 images, timestamps, TTS voice instructions, batch processing --- ### Pattern 8: Error Handling Always catch `OpenAI.APIError` and its subclasses. Re-throw unexpected errors. ```typescript try { const completion = await client.chat.completions.create({ model: "gpt-4o", messages: [{ role: "user", content: "Hello" }], }); } catch (error) { if (error instanceof OpenAI.APIError) { console.error( `API Error [${error.status}]: ${error.message} (${error.request_id})`, ); // Check subclasses: RateLimitError, AuthenticationError, BadRequestError, etc. } else { throw error; // Re-throw non-API errors } } ``` **Why good:** Specific error types with status codes, request ID for debugging, re-throws unexpected errors **See:** [examples/core.md](examples/core.md) for full production error handling, stream errors, error type hierarchy </patterns> --- <performance> ## Performance Optimization ### Model Selection for Cost/Speed ``` General purpose -> gpt-5.4 (most capable) or gpt-4o (proven, lower cost) Cost-sensitive / high-vol -> gpt-5-mini or gpt-5-nano (cheapest) Complex reasoning -> gpt-5.4 or o4-mini Structured output -> gpt-5.4 or gpt-4o (best schema adherence) Embeddings -> text-embedding-3-small (cheapest) or text-embedding-3-large (highest quality) Transcription -> whisper-1 or gpt-4o-transcribe (higher accuracy) TTS -> tts-1 (fast) or tts-1-hd (quality) or gpt-4o-mini-tts (voice control) Batch processing -> gpt-5-mini at 50% batch discount ``` ### Key Optimization Patterns - **Track token usage** via `completion.usage` for cost visibility - **Check `finish_reason === "length"`** to detect truncated output - **Use `temperature: 0`** for deterministic output (enables caching) - **Use `AbortController`** to cancel long-running requests - **Use Batch API** for high-volume jobs at 50% cost reduction </performance> --- <decision_framework> ## Decision Framework ### Which API to Use ``` Building a new application? +-- YES -> Need built-in tools (web search, file search, code interpreter)? | +-- YES -> Use Responses API (client.responses.create()) | +-- NO -> Need server-side conversation state? | +-- YES -> Use Responses API with store: true | +-- NO -> Either API works, prefer Responses for new code +-- Existing Chat Completions code? +-- Working fine? -> Keep using Chat Completions (fully supported) +-- Need new features? -> Consider migrating to Responses API ``` ### Which Model to Choose ``` What is your task? +-- General text generation -> gpt-5.4 (most capable) or gpt-4o (lower cost) +-- Fast + cheap simple tasks -> gpt-5-mini or gpt-5-nano +-- Complex reasoning / math -> gpt-5.4 or o4-mini +-- Structured output -> gpt-5.4 or gpt-4o (best schema adherence) +-- Vision (images) -> gpt-5.4 or gpt-4o +-- Embeddings -> text-embedding-3-small (default) or text-embedding-3-large +-- Transcription -> whisper-1 or gpt-4o-transcribe +-- Text-to-speech -> tts-1 (fast) or gpt-4o-mini-tts (voice instructions) +-- Batch processing -> gpt-5-mini (cheapest at 50% batch discount) ``` ### Streaming vs Non-Streaming ``` Is the response user-facing? +-- YES -> Use streaming (stream: true or .stream()) | +-- Need event-level control? -> .stream() with event handlers | +-- Simple text output? -> stream: true with for await +-- NO -> Use non-streaming +-- Background processing -> client.chat.completions.create() +-- Structured output -> client.chat.completions.parse() +-- High volume -> Batch API ``` ### When to Use This SDK vs a Provider-Agnostic SDK ``` Do you need multiple LLM providers (OpenAI + others)? +-- YES -> Not this skill's scope -- use a unified provider SDK +-- NO -> Do you need OpenAI-specific features? +-- YES -> Use OpenAI SDK directly | Examples: Responses API, Batch API, | Realtime API, built-in web search/file search +-- NO -> OpenAI SDK is simplest for OpenAI-only use ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Hardcoding API keys instead of using environment variables (security breach risk) - Using bare `catch` blocks without checking `OpenAI.APIError` (hides API errors) - Not consuming streams returned by `stream: true` (tokens are silently lost) - Using `JSON.parse()` on completion content without `zodResponseFormat` (fragile, no validation) - Sending full conversation history every request when Responses API's `previous_response_id` could manage state **Medium Priority Issues:** - Not setting `maxRetries` / `timeout` for production deployments (10 min default timeout may be too long) - Missing `developer` role message (no system instruction = unpredictable output style) - Using deprecated `system` role instead of `developer` role in Chat Completions - Not checking `finish_reason` for `'length'` truncation - Ignoring `usage` data (no cost visibility) **Common Mistakes:** - Confusing Responses API (`client.responses.create()`) with Chat Completions (`client.chat.completions.create()`) parameters -- they use different shapes - Using `messages` parameter with Responses API (it uses `input` and `instructions`) - Using `response_format` with models that don't support structured outputs (need gpt-4o or later) - Using `max_tokens` with reasoning models (o4-mini, gpt-5.x) -- use `max_completion_tokens` instead - Not handling the case where `completion.choices[0].message.tool_calls` is undefined - Forgetting that `runTools()` defaults to max 10 completions -- set `maxChatCompletions` explicitly **Gotchas & Edge Cases:** - The SDK auto-retries on 429 (rate limit) and 5xx errors -- 2 retries by default. Disable with `maxRetries: 0` if you handle retries yourself. - `stream: true` returns raw SSE chunks. Use `.stream()` helper for a nicer event-based API. - `client.chat.completions.parse()` throws `LengthFinishReasonError` if `finish_reason` is `'length'` and `ContentFilterFinishReasonError` if `'content_filter'`. - Embedding responses return `Array<number>` (the SDK requests base64 by default and decodes via Float32 internally for performance). No conversion needed -- you get a plain number array. - File uploads support `ReadStream`, `File`, `fetch()` Response, and `toFile()` helper -- use whichever matches your data source. - The Responses API's `store: true` enables server-side state but also means OpenAI stores your conversations. Set `store: false` for sensitive data. - `developer` role replaces `system` role in newer models (gpt-4o and later). - Batch API has a 24h completion window and 50,000 request limit per batch. - Audio transcription has a 25 MB file size limit. - Zod schemas with `zodResponseFormat` must use `additionalProperties: false` -- the SDK handles this automatically. - `zodTextFormat` and `zodResponseFormat` are NOT compatible with Zod v4 -- use Zod v3.x until the SDK adds v4 support. - The Assistants API is deprecated (sunset August 2026) -- use the Responses API for new code. </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST use the Responses API (`client.responses.create()`) for new projects -- it provides better performance, built-in tools, and server-side conversation state)** **(You MUST use `zodResponseFormat()` from `openai/helpers/zod` for structured outputs -- do NOT manually construct JSON schemas)** **(You MUST handle errors using `OpenAI.APIError` and its subclasses -- never use bare catch blocks without error type checking)** **(You MUST configure appropriate retries and timeouts for production use -- the SDK retries 2 times by default on 429/5xx errors)** **(You MUST never hardcode API keys -- always use environment variables via `process.env.OPENAI_API_KEY`)** **Failure to follow these rules will produce insecure, unreliable, or poorly-typed AI integrations.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.