Claude Skill

ai-provider-openai-sdk

Official OpenAI SDK patterns for TypeScript/Node.js — client setup, Chat Completions, Responses API, streaming, structured outputs, function calling, embeddings, vision, audio, and production best practices

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_ai-provider-openai-sdk_skills_ai-provider-openai-sdk-3a51ef5.zip · 22 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-provider-openai-sdk/skills/ai-provider-openai-sdk
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

OpenAI SDK Patterns

Quick Guide: Use the official openai npm package (v6+) to interact with OpenAI's API directly. Use client.responses.create() (Responses API) for new projects with built-in tools and server-side state, or client.chat.completions.create() (Chat Completions) for stateless chat flows. Use zodResponseFormat and client.chat.completions.parse() for structured outputs. Use .stream() or stream: true for streaming. Supports GPT-5.x family, GPT-4o, o4-mini, embeddings, vision, audio, and batch processing.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST use the Responses API (client.responses.create()) for new projects -- it provides better performance, built-in tools, and server-side conversation state)

(You MUST use zodResponseFormat() from openai/helpers/zod for structured outputs -- do NOT manually construct JSON schemas)

(You MUST handle errors using OpenAI.APIError and its subclasses -- never use bare catch blocks without error type checking)

(You MUST configure appropriate retries and timeouts for production use -- the SDK retries 2 times by default on 429/5xx errors)

(You MUST never hardcode API keys -- always use environment variables via process.env.OPENAI_API_KEY)

</critical_requirements>


Auto-detection: OpenAI, openai, client.chat.completions, client.responses.create, client.responses.parse, client.embeddings, client.audio, zodResponseFormat, zodTextFormat, zodFunction, zodResponsesFunction, runTools, GPT-5, GPT-4o, o4-mini, gpt-5-mini, text-embedding-3, whisper, tts, OPENAI_API_KEY, toFile

When to use:

  • Building applications that call OpenAI models directly (GPT-5.x, GPT-4o, o4-mini, etc.)
  • Implementing chat completions with streaming responses
  • Using the Responses API for agentic workflows with built-in tools (web search, file search, code interpreter)
  • Extracting structured data from LLM responses with Zod schema validation
  • Implementing function calling / tool use with the Chat Completions or Responses API
  • Creating embeddings for RAG pipelines or semantic search
  • Processing images with vision models or audio with Whisper/TTS
  • Running batch jobs for high-volume, cost-efficient processing

Key patterns covered:

  • Client initialization and configuration (retries, timeouts, proxies)
  • Chat Completions API (messages, streaming, function calling)
  • Responses API (input, instructions, built-in tools, server-side state)
  • Structured outputs with zodResponseFormat and client.chat.completions.parse()
  • Streaming with for await...of, .stream() helper, and event handling
  • Embeddings API (text-embedding-3-small, text-embedding-3-large)
  • Vision (image URLs, base64), Audio (Whisper transcription, TTS), Batch API
  • Error handling, retries, timeouts, and production best practices

When NOT to use:

  • Multi-provider applications where you need to switch between OpenAI, Anthropic, Google, etc. -- use a unified provider SDK instead
  • React-specific chat UI hooks (useChat, useCompletion) -- use a framework-integrated AI SDK
  • When you need a higher-level abstraction over multiple LLM providers

Examples Index





<decision_framework>

Decision Framework

Which API to Use

Building a new application?
+-- YES -> Need built-in tools (web search, file search, code interpreter)?
|   +-- YES -> Use Responses API (client.responses.create())
|   +-- NO -> Need server-side conversation state?
|       +-- YES -> Use Responses API with store: true
|       +-- NO -> Either API works, prefer Responses for new code
+-- Existing Chat Completions code?
    +-- Working fine? -> Keep using Chat Completions (fully supported)
    +-- Need new features? -> Consider migrating to Responses API

Which Model to Choose

What is your task?
+-- General text generation -> gpt-5.4 (most capable) or gpt-4o (lower cost)
+-- Fast + cheap simple tasks -> gpt-5-mini or gpt-5-nano
+-- Complex reasoning / math -> gpt-5.4 or o4-mini
+-- Structured output -> gpt-5.4 or gpt-4o (best schema adherence)
+-- Vision (images) -> gpt-5.4 or gpt-4o
+-- Embeddings -> text-embedding-3-small (default) or text-embedding-3-large
+-- Transcription -> whisper-1 or gpt-4o-transcribe
+-- Text-to-speech -> tts-1 (fast) or gpt-4o-mini-tts (voice instructions)
+-- Batch processing -> gpt-5-mini (cheapest at 50% batch discount)

Streaming vs Non-Streaming

Is the response user-facing?
+-- YES -> Use streaming (stream: true or .stream())
|   +-- Need event-level control? -> .stream() with event handlers
|   +-- Simple text output? -> stream: true with for await
+-- NO -> Use non-streaming
    +-- Background processing -> client.chat.completions.create()
    +-- Structured output -> client.chat.completions.parse()
    +-- High volume -> Batch API

When to Use This SDK vs a Provider-Agnostic SDK

Do you need multiple LLM providers (OpenAI + others)?
+-- YES -> Not this skill's scope -- use a unified provider SDK
+-- NO -> Do you need OpenAI-specific features?
    +-- YES -> Use OpenAI SDK directly
    |   Examples: Responses API, Batch API,
    |   Realtime API, built-in web search/file search
    +-- NO -> OpenAI SDK is simplest for OpenAI-only use

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Hardcoding API keys instead of using environment variables (security breach risk)
  • Using bare catch blocks without checking OpenAI.APIError (hides API errors)
  • Not consuming streams returned by stream: true (tokens are silently lost)
  • Using JSON.parse() on completion content without zodResponseFormat (fragile, no validation)
  • Sending full conversation history every request when Responses API's previous_response_id could manage state

Medium Priority Issues:

  • Not setting maxRetries / timeout for production deployments (10 min default timeout may be too long)
  • Missing developer role message (no system instruction = unpredictable output style)
  • Using deprecated system role instead of developer role in Chat Completions
  • Not checking finish_reason for 'length' truncation
  • Ignoring usage data (no cost visibility)

Common Mistakes:

  • Confusing Responses API (client.responses.create()) with Chat Completions (client.chat.completions.create()) parameters -- they use different shapes
  • Using messages parameter with Responses API (it uses input and instructions)
  • Using response_format with models that don't support structured outputs (need gpt-4o or later)
  • Using max_tokens with reasoning models (o4-mini, gpt-5.x) -- use max_completion_tokens instead
  • Not handling the case where completion.choices[0].message.tool_calls is undefined
  • Forgetting that runTools() defaults to max 10 completions -- set maxChatCompletions explicitly

Gotchas & Edge Cases:

  • The SDK auto-retries on 429 (rate limit) and 5xx errors -- 2 retries by default. Disable with maxRetries: 0 if you handle retries yourself.
  • stream: true returns raw SSE chunks. Use .stream() helper for a nicer event-based API.
  • client.chat.completions.parse() throws LengthFinishReasonError if finish_reason is 'length' and ContentFilterFinishReasonError if 'content_filter'.
  • Embedding responses return Array<number> (the SDK requests base64 by default and decodes via Float32 internally for performance). No conversion needed -- you get a plain number array.
  • File uploads support ReadStream, File, fetch() Response, and toFile() helper -- use whichever matches your data source.
  • The Responses API's store: true enables server-side state but also means OpenAI stores your conversations. Set store: false for sensitive data.
  • developer role replaces system role in newer models (gpt-4o and later).
  • Batch API has a 24h completion window and 50,000 request limit per batch.
  • Audio transcription has a 25 MB file size limit.
  • Zod schemas with zodResponseFormat must use additionalProperties: false -- the SDK handles this automatically.
  • zodTextFormat and zodResponseFormat are NOT compatible with Zod v4 -- use Zod v3.x until the SDK adds v4 support.
  • The Assistants API is deprecated (sunset August 2026) -- use the Responses API for new code.

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST use the Responses API (client.responses.create()) for new projects -- it provides better performance, built-in tools, and server-side conversation state)

(You MUST use zodResponseFormat() from openai/helpers/zod for structured outputs -- do NOT manually construct JSON schemas)

(You MUST handle errors using OpenAI.APIError and its subclasses -- never use bare catch blocks without error type checking)

(You MUST configure appropriate retries and timeouts for production use -- the SDK retries 2 times by default on 429/5xx errors)

(You MUST never hardcode API keys -- always use environment variables via process.env.OPENAI_API_KEY)

Failure to follow these rules will produce insecure, unreliable, or poorly-typed AI integrations.

</critical_reminders>

Files (skills)
  • examples
    • chat.md 3.1 KB
      # OpenAI SDK -- Chat Completions Examples
      
      > Chat Completions API patterns: basic completion, multi-turn conversations, system/developer messages, temperature, token control. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio
      
      ---
      
      ## Basic Chat Completion
      
      ```typescript
      // basic-chat.ts
      import OpenAI from "openai";
      
      const client = new OpenAI();
      
      async function chat(userMessage: string): Promise<string> {
        const completion = await client.chat.completions.create({
          model: "gpt-4o",
          messages: [
            {
              role: "developer",
              content: "You are a helpful assistant. Be concise.",
            },
            { role: "user", content: userMessage },
          ],
        });
      
        const content = completion.choices[0].message.content;
        if (!content) {
          throw new Error("No content in response");
        }
      
        console.log(`Tokens: ${completion.usage?.total_tokens}`);
        return content;
      }
      
      const answer = await chat("What is TypeScript in one sentence?");
      console.log(answer);
      ```
      
      ---
      
      ## Multi-Turn Conversations
      
      ```typescript
      import OpenAI from "openai";
      import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
      
      const client = new OpenAI();
      
      const messages: ChatCompletionMessageParam[] = [
        { role: "developer", content: "You are a TypeScript expert." },
        { role: "user", content: "What is a union type?" },
      ];
      
      const completion = await client.chat.completions.create({
        model: "gpt-4o",
        messages,
      });
      
      // Append assistant response for next turn
      const assistantMessage = completion.choices[0].message;
      messages.push(assistantMessage);
      messages.push({ role: "user", content: "Give me a real-world example." });
      
      const followUp = await client.chat.completions.create({
        model: "gpt-4o",
        messages,
      });
      ```
      
      ---
      
      ## Token Usage Tracking
      
      ```typescript
      const completion = await client.chat.completions.create({
        model: "gpt-4o",
        messages: [{ role: "user", content: "Hello" }],
      });
      
      const usage = completion.usage;
      if (usage) {
        console.log(`Prompt tokens: ${usage.prompt_tokens}`);
        console.log(`Completion tokens: ${usage.completion_tokens}`);
        console.log(`Total tokens: ${usage.total_tokens}`);
      
        // Cached tokens (prompt caching)
        if (usage.prompt_tokens_details?.cached_tokens) {
          console.log(`Cached: ${usage.prompt_tokens_details.cached_tokens}`);
        }
      }
      ```
      
      ---
      
      ## Controlling Output Length
      
      ```typescript
      const MAX_TOKENS = 500;
      
      const completion = await client.chat.completions.create({
        model: "gpt-4o",
        messages: [{ role: "user", content: "Summarize this article." }],
        max_tokens: MAX_TOKENS,
        temperature: 0, // Deterministic output for caching
      });
      
      if (completion.choices[0].finish_reason === "length") {
        console.warn("Output was truncated -- increase max_tokens");
      }
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • core.md 4.5 KB
      # OpenAI SDK -- Setup & Configuration Examples
      
      > Client initialization, environment config, production settings, Azure OpenAI, error handling, and abort patterns. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio
      
      ---
      
      ## Basic Client Setup
      
      ```typescript
      // lib/openai.ts
      import OpenAI from "openai";
      
      // Reads OPENAI_API_KEY from env automatically
      const client = new OpenAI();
      
      export { client };
      ```
      
      ---
      
      ## Production Configuration
      
      ```typescript
      // lib/openai.ts
      import OpenAI from "openai";
      
      const TIMEOUT_MS = 30_000;
      const MAX_RETRIES = 3;
      
      const client = new OpenAI({
        apiKey: process.env.OPENAI_API_KEY,
        timeout: TIMEOUT_MS,
        maxRetries: MAX_RETRIES,
      });
      
      export { client };
      ```
      
      ---
      
      ## Azure OpenAI
      
      ```typescript
      // lib/azure-openai.ts
      import { AzureOpenAI } from "openai";
      
      const azureClient = new AzureOpenAI({
        apiKey: process.env.AZURE_OPENAI_API_KEY,
        apiVersion: "2024-10-21",
        endpoint: process.env.AZURE_OPENAI_ENDPOINT,
      });
      
      export { azureClient };
      ```
      
      ---
      
      ## Production Error Handling
      
      ```typescript
      // error-handling.ts
      import OpenAI from "openai";
      
      const TIMEOUT_MS = 30_000;
      const MAX_RETRIES = 3;
      
      const client = new OpenAI({
        timeout: TIMEOUT_MS,
        maxRetries: MAX_RETRIES,
      });
      
      async function safeCompletion(prompt: string): Promise<string | null> {
        try {
          const completion = await client.chat.completions.create({
            model: "gpt-4o",
            messages: [
              { role: "developer", content: "You are a helpful assistant." },
              { role: "user", content: prompt },
            ],
          });
      
          // Check for truncation
          if (completion.choices[0].finish_reason === "length") {
            console.warn("Response was truncated");
          }
      
          return completion.choices[0].message.content;
        } catch (error) {
          if (error instanceof OpenAI.APIError) {
            console.error(`OpenAI API Error [${error.status}]: ${error.message}`);
            console.error(`Request ID: ${error.request_id}`);
      
            if (error instanceof OpenAI.RateLimitError) {
              console.error("Rate limited. SDK will auto-retry.");
              // If we get here, all retries were exhausted
              return null;
            }
      
            if (error instanceof OpenAI.AuthenticationError) {
              throw new Error(
                "Invalid API key. Check OPENAI_API_KEY environment variable.",
              );
            }
      
            if (error instanceof OpenAI.BadRequestError) {
              console.error("Invalid request parameters:", error.message);
              return null;
            }
      
            // Server errors (5xx) -- SDK auto-retries, if we're here all retries failed
            if (error instanceof OpenAI.InternalServerError) {
              console.error("OpenAI server error after all retries");
              return null;
            }
          }
      
          // Network/connection errors
          if (error instanceof OpenAI.APIConnectionError) {
            console.error("Network error:", error.message);
            return null;
          }
      
          // Unknown errors should be re-thrown
          throw error;
        }
      }
      
      const result = await safeCompletion("Hello!");
      if (result) {
        console.log(result);
      } else {
        console.error("Failed to get completion");
      }
      ```
      
      ---
      
      ## Stream Error Handling
      
      ```typescript
      const stream = client.chat.completions.stream({
        model: "gpt-4o",
        messages: [{ role: "user", content: "Hello" }],
      });
      
      stream.on("error", (error) => {
        if (error instanceof OpenAI.APIError) {
          console.error(`Stream API error: ${error.status}`);
        } else {
          console.error("Stream connection error:", error);
        }
      });
      
      stream.on("content", (delta) => {
        process.stdout.write(delta);
      });
      
      await stream.finalContent();
      ```
      
      ---
      
      ## Request Cancellation with AbortController
      
      ```typescript
      const controller = new AbortController();
      const ABORT_TIMEOUT_MS = 5_000;
      
      // Cancel after timeout
      setTimeout(() => controller.abort(), ABORT_TIMEOUT_MS);
      
      try {
        const completion = await client.chat.completions.create(
          { model: "gpt-4o", messages: [{ role: "user", content: "Hello" }] },
          { signal: controller.signal },
        );
      } catch (error) {
        if (error instanceof Error && error.name === "AbortError") {
          console.log("Request was cancelled");
        }
      }
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For per-request overrides, request ID tracking, and error type tables, see [reference.md](../reference.md)._
      
    • embeddings-vision-audio.md 9.7 KB
      # OpenAI SDK -- Embeddings, Vision & Audio Examples
      
      > Embeddings for semantic search, vision/image inputs, audio transcription and TTS, file uploads, and batch processing. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      
      ---
      
      ## Embeddings and Semantic Search
      
      ```typescript
      // embeddings.ts
      import OpenAI from "openai";
      
      const client = new OpenAI();
      const EMBEDDING_MODEL = "text-embedding-3-small";
      const SIMILARITY_THRESHOLD = 0.7;
      const TOP_K = 3;
      
      function cosineSimilarity(a: number[], b: number[]): number {
        let dot = 0;
        let normA = 0;
        let normB = 0;
        for (let i = 0; i < a.length; i++) {
          dot += a[i] * b[i];
          normA += a[i] * a[i];
          normB += b[i] * b[i];
        }
        return dot / (Math.sqrt(normA) * Math.sqrt(normB));
      }
      
      // Index documents
      const documents = [
        "TypeScript provides static type checking for JavaScript.",
        "React is a library for building user interfaces.",
        "Node.js is a JavaScript runtime built on V8.",
        "PostgreSQL is a powerful relational database.",
        "Docker containers package applications with dependencies.",
      ];
      
      const docEmbeddings = await client.embeddings.create({
        model: EMBEDDING_MODEL,
        input: documents,
      });
      
      const indexedDocs = documents.map((text, i) => ({
        text,
        embedding: Array.from(docEmbeddings.data[i].embedding),
      }));
      
      // Search
      async function search(
        query: string,
      ): Promise<Array<{ text: string; score: number }>> {
        const queryEmbedding = await client.embeddings.create({
          model: EMBEDDING_MODEL,
          input: query,
        });
      
        const queryVector = Array.from(queryEmbedding.data[0].embedding);
      
        return indexedDocs
          .map((doc) => ({
            text: doc.text,
            score: cosineSimilarity(queryVector, doc.embedding),
          }))
          .filter((r) => r.score > SIMILARITY_THRESHOLD)
          .sort((a, b) => b.score - a.score)
          .slice(0, TOP_K);
      }
      
      const results = await search("What is TypeScript?");
      results.forEach((r) => {
        console.log(`[${r.score.toFixed(3)}] ${r.text}`);
      });
      ```
      
      ---
      
      ## Reduced Embedding Dimensions
      
      ```typescript
      // Use fewer dimensions for cost/speed tradeoff
      const response = await client.embeddings.create({
        model: "text-embedding-3-large",
        input: "Some text to embed.",
        dimensions: 256, // Reduce from default 3072
      });
      ```
      
      ---
      
      ## Vision -- Image from URL
      
      ```typescript
      // vision.ts
      import OpenAI from "openai";
      
      const client = new OpenAI();
      
      async function analyzeImageUrl(
        imageUrl: string,
        question: string,
      ): Promise<string> {
        const response = await client.chat.completions.create({
          model: "gpt-4o",
          messages: [
            {
              role: "user",
              content: [
                { type: "text", text: question },
                { type: "image_url", image_url: { url: imageUrl } },
              ],
            },
          ],
        });
      
        return response.choices[0].message.content ?? "";
      }
      ```
      
      ---
      
      ## Vision -- Local Image (Base64)
      
      ```typescript
      import OpenAI from "openai";
      import { readFileSync } from "node:fs";
      
      const client = new OpenAI();
      
      async function analyzeLocalImage(
        imagePath: string,
        question: string,
      ): Promise<string> {
        const imageBuffer = readFileSync(imagePath);
        const base64Image = imageBuffer.toString("base64");
        const mimeType = imagePath.endsWith(".png") ? "image/png" : "image/jpeg";
      
        const response = await client.chat.completions.create({
          model: "gpt-4o",
          messages: [
            {
              role: "user",
              content: [
                { type: "text", text: question },
                {
                  type: "image_url",
                  image_url: {
                    url: `data:${mimeType};base64,${base64Image}`,
                    detail: "high", // 'low' | 'high' | 'auto'
                  },
                },
              ],
            },
          ],
        });
      
        return response.choices[0].message.content ?? "";
      }
      ```
      
      ---
      
      ## Vision -- Multiple Images
      
      ```typescript
      const response = await client.chat.completions.create({
        model: "gpt-4o",
        messages: [
          {
            role: "user",
            content: [
              { type: "text", text: "Compare these two images." },
              {
                type: "image_url",
                image_url: { url: "https://example.com/image1.jpg" },
              },
              {
                type: "image_url",
                image_url: { url: "https://example.com/image2.jpg" },
              },
            ],
          },
        ],
      });
      ```
      
      ---
      
      ## Audio Transcription (Speech-to-Text)
      
      ```typescript
      // audio.ts
      import OpenAI from "openai";
      import { createReadStream } from "node:fs";
      
      const client = new OpenAI();
      
      // Basic transcription
      async function transcribe(audioPath: string): Promise<string> {
        const transcription = await client.audio.transcriptions.create({
          model: "whisper-1",
          file: createReadStream(audioPath),
          language: "en",
        });
      
        return transcription.text;
      }
      
      // With word-level timestamps
      async function transcribeWithTimestamps(audioPath: string) {
        const transcription = await client.audio.transcriptions.create({
          model: "whisper-1",
          file: createReadStream(audioPath),
          response_format: "verbose_json",
          timestamp_granularities: ["word"],
        });
      
        return transcription.words?.map((w) => ({
          word: w.word,
          start: w.start,
          end: w.end,
        }));
      }
      
      const transcript = await transcribe("./recording.mp3");
      console.log("Transcript:", transcript);
      ```
      
      ---
      
      ## Text-to-Speech (TTS)
      
      ```typescript
      import OpenAI from "openai";
      import { writeFileSync } from "node:fs";
      
      const client = new OpenAI();
      
      // Basic TTS
      async function textToSpeech(
        text: string,
        outputPath: string,
        voice:
          | "alloy"
          | "ash"
          | "ballad"
          | "coral"
          | "echo"
          | "fable"
          | "nova"
          | "onyx"
          | "sage"
          | "shimmer" = "alloy",
      ): Promise<void> {
        const speech = await client.audio.speech.create({
          model: "tts-1",
          voice,
          input: text,
        });
      
        const buffer = Buffer.from(await speech.arrayBuffer());
        writeFileSync(outputPath, buffer);
        console.log(`Speech saved to ${outputPath}`);
      }
      
      // With voice instructions (gpt-4o-mini-tts only)
      async function textToSpeechWithStyle(
        text: string,
        outputPath: string,
        style: string,
      ): Promise<void> {
        const speech = await client.audio.speech.create({
          model: "gpt-4o-mini-tts",
          voice: "coral",
          input: text,
          instructions: style,
        });
      
        const buffer = Buffer.from(await speech.arrayBuffer());
        writeFileSync(outputPath, buffer);
      }
      
      await textToSpeech("Hello, welcome to the demo!", "./output.mp3", "nova");
      await textToSpeechWithStyle(
        "Breaking news from the tech world!",
        "./news.mp3",
        "Speak like an excited news anchor",
      );
      ```
      
      ---
      
      ## File Uploads
      
      ```typescript
      import OpenAI, { toFile } from "openai";
      import { createReadStream } from "node:fs";
      
      const client = new OpenAI();
      
      // From file path (Node.js ReadStream)
      const fileFromPath = await client.files.create({
        file: createReadStream("training-data.jsonl"),
        purpose: "fine-tune",
      });
      
      // From Buffer using toFile helper
      const fileFromBuffer = await client.files.create({
        file: await toFile(Buffer.from('{"prompt": "Hi"}'), "data.jsonl"),
        purpose: "fine-tune",
      });
      
      // From fetch Response
      const fileFromFetch = await client.files.create({
        file: await fetch("https://example.com/data.jsonl"),
        purpose: "fine-tune",
      });
      
      console.log(`Uploaded: ${fileFromPath.id}`);
      ```
      
      ---
      
      ## Batch Processing
      
      ```typescript
      // batch-processing.ts
      import OpenAI, { toFile } from "openai";
      
      const client = new OpenAI();
      const POLL_INTERVAL_MS = 30_000;
      
      interface BatchRequest {
        custom_id: string;
        method: "POST";
        url: string;
        body: {
          model: string;
          messages: Array<{ role: string; content: string }>;
        };
      }
      
      // Create batch input
      function createBatchInput(prompts: string[]): string {
        const requests: BatchRequest[] = prompts.map((prompt, index) => ({
          custom_id: `req-${index}`,
          method: "POST",
          url: "/v1/chat/completions",
          body: {
            model: "gpt-4o-mini",
            messages: [
              {
                role: "developer",
                content: "Classify the sentiment as positive, negative, or neutral.",
              },
              { role: "user", content: prompt },
            ],
          },
        }));
      
        return requests.map((r) => JSON.stringify(r)).join("\n");
      }
      
      async function runBatch(prompts: string[]): Promise<string> {
        // Upload input file
        const jsonl = createBatchInput(prompts);
        const inputFile = await client.files.create({
          file: await toFile(Buffer.from(jsonl), "batch-input.jsonl"),
          purpose: "batch",
        });
      
        // Create batch
        const batch = await client.batches.create({
          input_file_id: inputFile.id,
          endpoint: "/v1/chat/completions",
          completion_window: "24h",
        });
      
        console.log(`Batch ${batch.id} created. Polling...`);
      
        // Poll for completion
        let status = batch;
        while (
          !["completed", "failed", "cancelled", "expired"].includes(status.status)
        ) {
          await new Promise((resolve) => setTimeout(resolve, POLL_INTERVAL_MS));
          status = await client.batches.retrieve(batch.id);
          console.log(
            `Status: ${status.status} (${status.request_counts?.completed ?? 0}/${status.request_counts?.total ?? 0})`,
          );
        }
      
        if (status.status !== "completed" || !status.output_file_id) {
          throw new Error(`Batch ${status.status}: ${JSON.stringify(status.errors)}`);
        }
      
        // Download results
        const outputFile = await client.files.content(status.output_file_id);
        return outputFile.text();
      }
      
      const prompts = [
        "I love this product! It exceeded my expectations.",
        "The service was terrible and the food was cold.",
        "The meeting was rescheduled to next Tuesday.",
      ];
      
      const results = await runBatch(prompts);
      console.log("Batch results:", results);
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • openai-sdk.md 552 B
      # Deprecated -- Examples Have Been Split
      
      This file has been replaced by topic-specific example files:
      
      - [core.md](core.md) -- Client setup, error handling, configuration
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio, files, batch
      
      See [SKILL.md](../SKILL.md) for the examples index.
      
    • streaming.md 2.9 KB
      # OpenAI SDK -- Streaming Examples
      
      > Streaming patterns for Chat Completions and Responses API: `stream: true` with async iterators, `.stream()` event-based helper, Responses API streaming. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat Completions API
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio
      
      ---
      
      ## Basic Streaming with `for await`
      
      ```typescript
      // streaming-chat.ts
      import OpenAI from "openai";
      
      const client = new OpenAI();
      
      // Chat Completions streaming
      const stream = await client.chat.completions.create({
        model: "gpt-4o",
        messages: [
          { role: "developer", content: "You are a helpful assistant." },
          { role: "user", content: "Explain async/await in TypeScript." },
        ],
        stream: true,
      });
      
      for await (const chunk of stream) {
        const content = chunk.choices[0]?.delta?.content;
        if (content) {
          process.stdout.write(content);
        }
      }
      console.log(); // newline
      ```
      
      ---
      
      ## Event-Based Streaming with `.stream()` Helper
      
      ```typescript
      import OpenAI from "openai";
      
      const client = new OpenAI();
      
      async function streamWithEvents(prompt: string): Promise<string> {
        const stream = client.chat.completions.stream({
          model: "gpt-4o",
          messages: [
            { role: "developer", content: "You are a helpful assistant." },
            { role: "user", content: prompt },
          ],
        });
      
        stream.on("content", (delta) => {
          process.stdout.write(delta);
        });
      
        stream.on("error", (error) => {
          console.error("Stream error:", error);
        });
      
        const content = await stream.finalContent();
        console.log(); // newline
        return content ?? "";
      }
      
      const result = await streamWithEvents("Explain promises in JavaScript.");
      console.log("Final result length:", result.length);
      ```
      
      ---
      
      ## Responses API Streaming
      
      ```typescript
      import OpenAI from "openai";
      
      const client = new OpenAI();
      
      const stream = await client.responses.create({
        model: "gpt-4o",
        input: "Write a haiku about TypeScript.",
        stream: true,
      });
      
      for await (const event of stream) {
        if (event.type === "response.output_text.delta") {
          process.stdout.write(event.delta);
        }
      }
      console.log();
      ```
      
      ---
      
      ## Stream Abort
      
      ```typescript
      const stream = client.chat.completions.stream({
        model: "gpt-4o",
        messages: [{ role: "user", content: "Tell me a long story." }],
      });
      
      // Cancel after a timeout
      const ABORT_TIMEOUT_MS = 5_000;
      setTimeout(() => stream.abort(), ABORT_TIMEOUT_MS);
      
      stream.on("content", (delta) => {
        process.stdout.write(delta);
      });
      
      try {
        await stream.finalContent();
      } catch (error) {
        console.log("Stream aborted");
      }
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For stream method signatures, event types, and API tables, see [reference.md](../reference.md)._
      
    • structured-output.md 4.4 KB
      # OpenAI SDK -- Structured Output Examples
      
      > Type-safe structured responses with Zod: `zodResponseFormat` for Chat Completions, `zodTextFormat` for Responses API, refusal handling, complex schemas. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio
      
      ---
      
      ## Chat Completions Structured Output with `zodResponseFormat`
      
      ```typescript
      // structured-output.ts
      import OpenAI from "openai";
      import { zodResponseFormat } from "openai/helpers/zod";
      import { z } from "zod";
      
      const client = new OpenAI();
      
      // Define the schema for extracted data
      const ArticleSummary = z.object({
        title: z.string(),
        summary: z.string(),
        keyPoints: z.array(z.string()),
        sentiment: z.enum(["positive", "negative", "neutral"]),
        wordCount: z.number(),
      });
      
      type ArticleSummary = z.infer<typeof ArticleSummary>;
      
      async function extractArticleSummary(
        articleText: string,
      ): Promise<ArticleSummary | null> {
        const completion = await client.chat.completions.parse({
          model: "gpt-4o",
          messages: [
            {
              role: "developer",
              content: "Extract a structured summary from the provided article text.",
            },
            { role: "user", content: articleText },
          ],
          response_format: zodResponseFormat(ArticleSummary, "article_summary"),
        });
      
        const message = completion.choices[0].message;
      
        // Handle safety refusals
        if (message.refusal) {
          console.warn("Model refused:", message.refusal);
          return null;
        }
      
        return message.parsed;
      }
      
      const article = `
      TypeScript 5.8 brings exciting new features including improved type inference,
      better error messages, and performance optimizations. The release focuses on
      developer experience improvements that make everyday coding more productive.
      `;
      
      const summary = await extractArticleSummary(article);
      if (summary) {
        console.log(`Title: ${summary.title}`);
        console.log(`Sentiment: ${summary.sentiment}`);
        console.log("Key Points:");
        summary.keyPoints.forEach((point) => console.log(`  - ${point}`));
      }
      ```
      
      ---
      
      ## Handling Refusals
      
      ```typescript
      import OpenAI from "openai";
      import { zodResponseFormat } from "openai/helpers/zod";
      import { z } from "zod";
      
      const client = new OpenAI();
      
      const SomeSchema = z.object({
        content: z.string(),
      });
      
      const completion = await client.chat.completions.parse({
        model: "gpt-4o",
        messages: [{ role: "user", content: "Generate harmful content" }],
        response_format: zodResponseFormat(SomeSchema, "output"),
      });
      
      const message = completion.choices[0].message;
      
      if (message.refusal) {
        console.log("Model refused:", message.refusal);
      } else if (message.parsed) {
        console.log("Parsed output:", message.parsed);
      }
      ```
      
      ---
      
      ## Responses API Structured Output with `zodTextFormat`
      
      ```typescript
      import OpenAI from "openai";
      import { zodTextFormat } from "openai/helpers/zod";
      import { z } from "zod";
      
      const client = new OpenAI();
      
      const PersonSchema = z.object({
        name: z.string(),
        age: z.number(),
      });
      
      const response = await client.responses.parse({
        model: "gpt-4o",
        input: "Jane is 54 years old.",
        text: {
          format: zodTextFormat(PersonSchema, "person"),
        },
      });
      
      console.log(response.output_parsed);
      // { name: 'Jane', age: 54 }
      ```
      
      ---
      
      ## Manual JSON Schema (Anti-Pattern)
      
      Avoid this -- use `zodResponseFormat` instead:
      
      ```typescript
      // BAD: Manually constructing JSON schema
      const completion = await client.chat.completions.create({
        model: "gpt-4o",
        messages: [{ role: "user", content: "Extract data" }],
        response_format: {
          type: "json_schema",
          json_schema: {
            name: "event",
            strict: true,
            schema: {
              type: "object",
              properties: {
                name: { type: "string" },
                date: { type: "string" },
              },
              required: ["name", "date"],
              additionalProperties: false,
            },
          },
        },
      });
      // Then manually JSON.parse the content -- error-prone
      const data = JSON.parse(completion.choices[0].message.content ?? "{}");
      ```
      
      **Why bad:** Manual JSON schema is verbose and error-prone, no type safety, manual parsing can fail. Use `zodResponseFormat` + `.parse()` instead.
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • tools.md 7.5 KB
      # OpenAI SDK -- Tool/Function Calling Examples
      
      > Function calling patterns: manual tool definitions, Zod-based tools with `zodFunction`, automated tool loops with `runTools`, Responses API function calling. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [embeddings-vision-audio.md](embeddings-vision-audio.md) -- Embeddings, vision, audio
      
      ---
      
      ## Chat Completions with Manual Tool Definitions
      
      ```typescript
      import OpenAI from "openai";
      import type { ChatCompletionTool } from "openai/resources/chat/completions";
      
      const client = new OpenAI();
      
      const tools: ChatCompletionTool[] = [
        {
          type: "function",
          function: {
            name: "get_weather",
            description: "Get the current weather for a location",
            parameters: {
              type: "object",
              properties: {
                location: { type: "string", description: "City name" },
                unit: { type: "string", enum: ["celsius", "fahrenheit"] },
              },
              required: ["location"],
              additionalProperties: false,
            },
            strict: true,
          },
        },
      ];
      
      const completion = await client.chat.completions.create({
        model: "gpt-4o",
        messages: [{ role: "user", content: "What is the weather in Tokyo?" }],
        tools,
      });
      
      const toolCall = completion.choices[0].message.tool_calls?.[0];
      if (toolCall) {
        const args = JSON.parse(toolCall.function.arguments);
        console.log(`Call ${toolCall.function.name} with:`, args);
      }
      ```
      
      ---
      
      ## Zod-Based Tools with `zodFunction`
      
      ```typescript
      // function-calling.ts
      import OpenAI from "openai";
      import { zodFunction } from "openai/helpers/zod";
      import { z } from "zod";
      
      const client = new OpenAI();
      
      // Define tool schemas with Zod
      const GetWeatherParams = z.object({
        location: z.string().describe('City name, e.g. "San Francisco"'),
        unit: z.enum(["celsius", "fahrenheit"]).default("celsius"),
      });
      
      const SearchDatabaseParams = z.object({
        query: z.string().describe("Search query string"),
        limit: z.number().default(10).describe("Max results to return"),
      });
      
      // Tool implementations
      async function getWeather(
        args: z.infer<typeof GetWeatherParams>,
      ): Promise<string> {
        // In production, call a real weather API
        return JSON.stringify({
          location: args.location,
          temperature: 22,
          unit: args.unit,
          condition: "sunny",
        });
      }
      
      async function searchDatabase(
        args: z.infer<typeof SearchDatabaseParams>,
      ): Promise<string> {
        // In production, query your database
        return JSON.stringify({
          results: [{ id: 1, title: `Result for: ${args.query}` }],
          total: 1,
        });
      }
      
      // Map tool names to implementations
      const toolImplementations: Record<string, (args: unknown) => Promise<string>> =
        {
          get_weather: getWeather as (args: unknown) => Promise<string>,
          search_database: searchDatabase as (args: unknown) => Promise<string>,
        };
      
      // Create completion with tools
      const completion = await client.chat.completions.parse({
        model: "gpt-4o",
        messages: [
          {
            role: "developer",
            content: "You help users by calling available tools.",
          },
          { role: "user", content: "What is the weather in Tokyo?" },
        ],
        tools: [
          zodFunction({ name: "get_weather", parameters: GetWeatherParams }),
          zodFunction({ name: "search_database", parameters: SearchDatabaseParams }),
        ],
      });
      
      // Process tool calls
      const message = completion.choices[0].message;
      if (message.tool_calls && message.tool_calls.length > 0) {
        for (const toolCall of message.tool_calls) {
          const fnName = toolCall.function.name;
          const args = JSON.parse(toolCall.function.arguments);
          console.log(`Calling ${fnName} with:`, args);
      
          const impl = toolImplementations[fnName];
          if (impl) {
            const result = await impl(args);
            console.log(`Result: ${result}`);
          }
        }
      }
      ```
      
      ---
      
      ## Automated Tool Loop with `runTools`
      
      ```typescript
      // run-tools.ts
      import OpenAI from "openai";
      
      const client = new OpenAI();
      
      const MAX_TOOL_CALLS = 5;
      
      async function getWeather(args: { location: string }): Promise<string> {
        return `Weather in ${args.location}: 22C, sunny`;
      }
      
      async function getTime(args: { timezone: string }): Promise<string> {
        return `Time in ${args.timezone}: ${new Date().toLocaleTimeString()}`;
      }
      
      const runner = client.chat.completions.runTools({
        model: "gpt-4o",
        messages: [
          { role: "developer", content: "Help users with weather and time queries." },
          {
            role: "user",
            content: "What is the weather and current time in London?",
          },
        ],
        tools: [
          {
            type: "function",
            function: {
              function: getWeather,
              parse: JSON.parse,
              description: "Get current weather for a location",
              parameters: {
                type: "object",
                properties: {
                  location: { type: "string", description: "City name" },
                },
                required: ["location"],
              },
            },
          },
          {
            type: "function",
            function: {
              function: getTime,
              parse: JSON.parse,
              description: "Get current time in a timezone",
              parameters: {
                type: "object",
                properties: {
                  timezone: {
                    type: "string",
                    description: "Timezone, e.g. Europe/London",
                  },
                },
                required: ["timezone"],
              },
            },
          },
        ],
        maxChatCompletions: MAX_TOOL_CALLS,
      });
      
      // Monitor the tool execution loop
      runner.on("message", (msg) => {
        if (msg.role === "assistant" && msg.tool_calls) {
          console.log(`[Tool calls: ${msg.tool_calls.length}]`);
        }
        if (msg.role === "tool") {
          console.log(`[Tool result received]`);
        }
      });
      
      const finalContent = await runner.finalContent();
      console.log("\nFinal answer:", finalContent);
      ```
      
      ---
      
      ## Responses API Function Calling
      
      ```typescript
      // responses-function-calling.ts
      import OpenAI from "openai";
      
      const client = new OpenAI();
      
      // Define the function tool for Responses API
      const response = await client.responses.create({
        model: "gpt-4o",
        instructions: "You are a helpful assistant with access to weather data.",
        input: "What is the weather like in San Francisco and New York?",
        tools: [
          {
            type: "function",
            name: "get_weather",
            description: "Get current weather for a city",
            parameters: {
              type: "object",
              properties: {
                location: { type: "string", description: "City name" },
              },
              required: ["location"],
              additionalProperties: false,
            },
          },
        ],
      });
      
      // Process function call outputs
      const functionCalls = response.output.filter(
        (item) => item.type === "function_call",
      );
      
      for (const call of functionCalls) {
        console.log(`Function: ${call.name}`);
        console.log(`Arguments: ${call.arguments}`);
        console.log(`Call ID: ${call.call_id}`);
      }
      
      // Submit function results back to continue the conversation
      if (functionCalls.length > 0) {
        const toolOutputs = functionCalls.map((call) => ({
          type: "function_call_output" as const,
          call_id: call.call_id,
          output: JSON.stringify({ temperature: 22, condition: "sunny" }),
        }));
      
        const followUp = await client.responses.create({
          model: "gpt-4o",
          input: toolOutputs,
          previous_response_id: response.id,
          store: true,
        });
      
        console.log("Final answer:", followUp.output_text);
      }
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
  • reference.md 12.9 KB
    # OpenAI SDK Quick Reference
    
    > Client configuration, model IDs, API methods, error types, and helper functions. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Package Installation
    
    ```bash
    # Core package (always required)
    npm install openai
    
    # For structured outputs (optional but recommended)
    npm install zod
    ```
    
    ---
    
    ## Client Configuration
    
    ```typescript
    import OpenAI from "openai";
    
    const client = new OpenAI({
      apiKey: process.env.OPENAI_API_KEY, // Auto-reads from env if not set
      timeout: 30_000, // Request timeout in ms (default: 600_000 = 10 min)
      maxRetries: 3, // Retry count on 429/5xx (default: 2)
      baseURL: "https://api.openai.com/v1", // Override for proxies
      dangerouslyAllowBrowser: false, // Must be true for browser usage
    });
    ```
    
    ### Environment Variables
    
    | Variable          | Purpose                    |
    | ----------------- | -------------------------- |
    | `OPENAI_API_KEY`  | API key (auto-detected)    |
    | `OPENAI_ORG_ID`   | Organization ID (optional) |
    | `OPENAI_BASE_URL` | Custom base URL (optional) |
    
    ---
    
    ## Model IDs
    
    ### Language Models (Chat / Text Generation)
    
    | Model         | Use Case                                  | Context Window | Notes                                  |
    | ------------- | ----------------------------------------- | -------------- | -------------------------------------- |
    | `gpt-5.4`     | Most capable, general purpose + reasoning | 1M             | Recommended for new projects           |
    | `gpt-5-mini`  | Cost-optimized, balanced speed/quality    | 1M             | Replaces gpt-4o-mini for most uses     |
    | `gpt-5-nano`  | High-throughput, straightforward tasks    | 1M             | Cheapest GPT-5 variant                 |
    | `gpt-4o`      | General purpose, proven                   | 128K           | Still available, lower cost than GPT-5 |
    | `gpt-4o-mini` | Fast, cheap, simple tasks                 | 128K           | Still available, migrate to gpt-5-mini |
    | `o4-mini`     | Fast reasoning                            | 200K           | Use `max_completion_tokens`            |
    
    ### Embedding Models
    
    | Model                    | Dimensions (default) | Max Input   |
    | ------------------------ | -------------------- | ----------- |
    | `text-embedding-3-small` | 1536                 | 8191 tokens |
    | `text-embedding-3-large` | 3072                 | 8191 tokens |
    
    ### Audio Models
    
    | Model                    | Type          | Use Case                    |
    | ------------------------ | ------------- | --------------------------- |
    | `whisper-1`              | Transcription | General speech-to-text      |
    | `gpt-4o-mini-transcribe` | Transcription | Fewer hallucinations        |
    | `gpt-4o-transcribe`      | Transcription | Higher accuracy             |
    | `tts-1`                  | TTS           | Fast text-to-speech         |
    | `tts-1-hd`               | TTS           | High-quality text-to-speech |
    | `gpt-4o-mini-tts`        | TTS           | Voice instruction support   |
    
    ### TTS Voices
    
    `alloy`, `ash`, `ballad`, `coral`, `echo`, `fable`, `nova`, `onyx`, `sage`, `shimmer`
    
    ---
    
    ## API Methods Reference
    
    ### Chat Completions API
    
    ```typescript
    // Standard completion
    const completion = await client.chat.completions.create({
      model: "gpt-4o", // Required
      messages: [], // Required: ChatCompletionMessageParam[]
      temperature: 0.7, // 0-2 (default: 1)
      max_tokens: 1000, // Max output tokens (use max_completion_tokens for reasoning models)
      top_p: 1, // Nucleus sampling
      frequency_penalty: 0, // -2 to 2
      presence_penalty: 0, // -2 to 2
      tools: [], // ChatCompletionTool[]
      tool_choice: "auto", // 'auto' | 'required' | 'none' | { type: 'function', function: { name: string } }
      response_format: undefined, // zodResponseFormat() or { type: 'json_object' }
      stream: false, // Enable streaming
      stop: undefined, // string | string[] -- stop sequences
      seed: undefined, // number -- for reproducibility
      logprobs: false, // Include log probabilities
      user: undefined, // End-user ID for abuse monitoring
    });
    
    // Structured output parsing
    const parsed = await client.chat.completions.parse({
      model: "gpt-4o",
      messages: [],
      response_format: zodResponseFormat(schema, "name"),
    });
    
    // Event-based streaming
    const stream = client.chat.completions.stream({
      model: "gpt-4o",
      messages: [],
    });
    
    // Automated tool execution
    const runner = client.chat.completions.runTools({
      model: "gpt-4o",
      messages: [],
      tools: [],
      maxChatCompletions: 10, // Max tool call loops (default: 10)
    });
    ```
    
    ### Responses API
    
    ```typescript
    const response = await client.responses.create({
      model: "gpt-4o", // Required
      input: "", // Required: string or InputItem[]
      instructions: "", // System instruction (replaces 'developer' role)
      tools: [], // Built-in or function tools
      previous_response_id: "", // Chain conversations
      store: true, // Enable server-side state
      stream: false, // Enable streaming
      text: {
        // Structured output config
        format: zodTextFormat(schema, "name"),
      },
    });
    
    // Access helpers
    response.output_text; // Direct text access
    response.output; // Array of typed output Items
    response.id; // Response ID for chaining
    ```
    
    ### Embeddings
    
    ```typescript
    const response = await client.embeddings.create({
      model: "text-embedding-3-small", // Required
      input: "", // string or string[]
      dimensions: undefined, // Reduce dimensions (optional)
      encoding_format: "float", // 'float' | 'base64'
    });
    ```
    
    ### Audio
    
    ```typescript
    // Transcription
    const transcription = await client.audio.transcriptions.create({
      model: "whisper-1", // Required
      file: readStream, // Required: ReadStream or File
      language: "en", // ISO 639-1 code
      response_format: "json", // 'json' | 'text' | 'srt' | 'vtt' | 'verbose_json'
      temperature: 0, // 0-1
      timestamp_granularities: [], // ['word', 'segment'] (verbose_json only)
    });
    
    // Text-to-Speech
    const speech = await client.audio.speech.create({
      model: "tts-1", // Required
      voice: "alloy", // Required
      input: "", // Required: text to speak
      response_format: "mp3", // mp3, opus, aac, flac, wav, pcm
      speed: 1.0, // 0.25-4.0
      instructions: "", // Voice style (gpt-4o-mini-tts only)
    });
    
    // Translation (to English)
    const translation = await client.audio.translations.create({
      model: "whisper-1",
      file: readStream,
    });
    ```
    
    ### Files
    
    ```typescript
    // Upload
    const file = await client.files.create({
      file: readStream, // ReadStream | File | Response | toFile()
      purpose: "fine-tune", // 'fine-tune' | 'batch'
    });
    
    // List
    const files = await client.files.list();
    
    // Retrieve
    const fileInfo = await client.files.retrieve("file-abc123");
    
    // Download content
    const content = await client.files.content("file-abc123");
    
    // Delete
    await client.files.del("file-abc123");
    ```
    
    ### Batch API
    
    ```typescript
    // Create batch
    const batch = await client.batches.create({
      input_file_id: "file-abc123",
      endpoint: "/v1/chat/completions", // or '/v1/embeddings'
      completion_window: "24h",
      metadata: {},
    });
    
    // Retrieve status
    const status = await client.batches.retrieve("batch-abc123");
    
    // List batches
    const batches = await client.batches.list();
    
    // Cancel batch
    await client.batches.cancel("batch-abc123");
    ```
    
    ---
    
    ## Helper Functions
    
    ```typescript
    // Structured output with Zod
    import { zodResponseFormat } from "openai/helpers/zod";
    import { zodTextFormat } from "openai/helpers/zod";
    import { zodFunction } from "openai/helpers/zod";
    import { zodResponsesFunction } from "openai/helpers/zod";
    
    // Chat Completions structured output
    zodResponseFormat(zodSchema, "schema_name");
    
    // Responses API structured output
    zodTextFormat(zodSchema, "schema_name");
    
    // Chat Completions function tool from Zod schema
    zodFunction({ name: "tool_name", parameters: zodSchema });
    
    // Responses API function tool from Zod schema
    zodResponsesFunction({ name: "tool_name", parameters: zodSchema });
    
    // File conversion helper
    import { toFile } from "openai";
    await toFile(buffer, "filename.ext");
    ```
    
    ---
    
    ## Error Types
    
    | Error Class                 | HTTP Status | Auto-Retried? |
    | --------------------------- | ----------- | ------------- |
    | `BadRequestError`           | 400         | No            |
    | `AuthenticationError`       | 401         | No            |
    | `PermissionDeniedError`     | 403         | No            |
    | `NotFoundError`             | 404         | No            |
    | `UnprocessableEntityError`  | 422         | No            |
    | `RateLimitError`            | 429         | Yes           |
    | `InternalServerError`       | >= 500      | Yes           |
    | `APIConnectionError`        | N/A         | Yes           |
    | `APIConnectionTimeoutError` | N/A         | Yes           |
    
    All errors extend `OpenAI.APIError` with properties:
    
    - `.status` -- HTTP status code
    - `.message` -- Error message
    - `.request_id` -- Request ID for debugging
    - `.headers` -- Response headers
    
    ---
    
    ## Streaming Events (.stream() Helper)
    
    | Event                                 | Arguments                                                         | Description                      |
    | ------------------------------------- | ----------------------------------------------------------------- | -------------------------------- |
    | `connect`                             | `()`                                                              | Connection established           |
    | `chunk`                               | `(chunk, snapshot)`                                               | Raw API chunk received           |
    | `content`                             | `(delta, snapshot)`                                               | Text content delta               |
    | `content.delta`                       | `({ delta, snapshot, parsed })`                                   | Content chunk with full snapshot |
    | `content.done`                        | `({ content, parsed })`                                           | Content generation complete      |
    | `refusal.delta`                       | `({ delta, snapshot })`                                           | Refusal content delta            |
    | `refusal.done`                        | `({ refusal })`                                                   | Refusal content complete         |
    | `tool_calls.function.arguments.delta` | `({ name, index, arguments, arguments_delta, parsed_arguments })` | Tool argument streaming          |
    | `tool_calls.function.arguments.done`  | `({ name, index, arguments, parsed_arguments })`                  | Tool arguments complete          |
    | `error`                               | `(error)`                                                         | Stream error                     |
    | `end`                                 | `()`                                                              | Stream finished                  |
    
    ### Stream Methods
    
    ```typescript
    await stream.finalContent(); // Promise<string> -- last assistant content
    await stream.finalMessage(); // Promise<ChatCompletionMessage>
    await stream.allChatCompletions(); // Promise<ChatCompletion[]>
    stream.abort(); // Cancel stream and network request
    stream.controller; // Underlying AbortController
    ```
    
    ---
    
    ## Responses API Event Types (Streaming)
    
    | Event Type                               | Description                 |
    | ---------------------------------------- | --------------------------- |
    | `response.created`                       | Response object created     |
    | `response.output_text.delta`             | Text content delta          |
    | `response.output_text.done`              | Text generation complete    |
    | `response.function_call_arguments.delta` | Function argument streaming |
    | `response.function_call_arguments.done`  | Function arguments complete |
    | `response.completed`                     | Full response complete      |
    | `error`                                  | Error occurred              |
    
    ---
    
    ## Message Roles
    
    | Role        | API              | Description                                            |
    | ----------- | ---------------- | ------------------------------------------------------ |
    | `developer` | Chat Completions | System instruction (replaces `system` in newer models) |
    | `system`    | Chat Completions | Legacy system instruction                              |
    | `user`      | Both             | User input                                             |
    | `assistant` | Chat Completions | Model response                                         |
    | `tool`      | Chat Completions | Tool result                                            |
    
    ---
    
    ## Per-Request Overrides
    
    ```typescript
    // Override retries and timeout for a single request
    await client.chat.completions.create(
      { model: 'gpt-4o', messages: [...] },
      {
        maxRetries: 5,
        timeout: 60_000,
        signal: abortController.signal,
        headers: { 'X-Custom-Header': 'value' },
      },
    );
    ```
    
    ---
    
    ## Request ID Tracking
    
    ```typescript
    // From response object
    completion._request_id;
    
    // From raw response headers
    const { data, response } = await client.chat.completions
      .create({ model: 'gpt-4o', messages: [...] })
      .withResponse();
    
    response.headers.get('x-request-id');
    ```
    
  • SKILL.md 20.6 KB
    ---
    name: ai-provider-openai-sdk
    description: Official OpenAI SDK patterns for TypeScript/Node.js — client setup, Chat Completions, Responses API, streaming, structured outputs, function calling, embeddings, vision, audio, and production best practices
    ---
    
    # OpenAI SDK Patterns
    
    > **Quick Guide:** Use the official `openai` npm package (v6+) to interact with OpenAI's API directly. Use `client.responses.create()` (Responses API) for new projects with built-in tools and server-side state, or `client.chat.completions.create()` (Chat Completions) for stateless chat flows. Use `zodResponseFormat` and `client.chat.completions.parse()` for structured outputs. Use `.stream()` or `stream: true` for streaming. Supports GPT-5.x family, GPT-4o, o4-mini, embeddings, vision, audio, and batch processing.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST use the Responses API (`client.responses.create()`) for new projects -- it provides better performance, built-in tools, and server-side conversation state)**
    
    **(You MUST use `zodResponseFormat()` from `openai/helpers/zod` for structured outputs -- do NOT manually construct JSON schemas)**
    
    **(You MUST handle errors using `OpenAI.APIError` and its subclasses -- never use bare catch blocks without error type checking)**
    
    **(You MUST configure appropriate retries and timeouts for production use -- the SDK retries 2 times by default on 429/5xx errors)**
    
    **(You MUST never hardcode API keys -- always use environment variables via `process.env.OPENAI_API_KEY`)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** OpenAI, openai, client.chat.completions, client.responses.create, client.responses.parse, client.embeddings, client.audio, zodResponseFormat, zodTextFormat, zodFunction, zodResponsesFunction, runTools, GPT-5, GPT-4o, o4-mini, gpt-5-mini, text-embedding-3, whisper, tts, OPENAI_API_KEY, toFile
    
    **When to use:**
    
    - Building applications that call OpenAI models directly (GPT-5.x, GPT-4o, o4-mini, etc.)
    - Implementing chat completions with streaming responses
    - Using the Responses API for agentic workflows with built-in tools (web search, file search, code interpreter)
    - Extracting structured data from LLM responses with Zod schema validation
    - Implementing function calling / tool use with the Chat Completions or Responses API
    - Creating embeddings for RAG pipelines or semantic search
    - Processing images with vision models or audio with Whisper/TTS
    - Running batch jobs for high-volume, cost-efficient processing
    
    **Key patterns covered:**
    
    - Client initialization and configuration (retries, timeouts, proxies)
    - Chat Completions API (messages, streaming, function calling)
    - Responses API (input, instructions, built-in tools, server-side state)
    - Structured outputs with `zodResponseFormat` and `client.chat.completions.parse()`
    - Streaming with `for await...of`, `.stream()` helper, and event handling
    - Embeddings API (`text-embedding-3-small`, `text-embedding-3-large`)
    - Vision (image URLs, base64), Audio (Whisper transcription, TTS), Batch API
    - Error handling, retries, timeouts, and production best practices
    
    **When NOT to use:**
    
    - Multi-provider applications where you need to switch between OpenAI, Anthropic, Google, etc. -- use a unified provider SDK instead
    - React-specific chat UI hooks (`useChat`, `useCompletion`) -- use a framework-integrated AI SDK
    - When you need a higher-level abstraction over multiple LLM providers
    
    ---
    
    ## Examples Index
    
    - [Core: Setup & Configuration](examples/core.md) -- Client init, production config, Azure, error handling, request overrides
    - [Chat Completions](examples/chat.md) -- Basic chat, multi-turn, token tracking, output length control
    - [Streaming](examples/streaming.md) -- `stream: true`, `.stream()` helper, Responses API streaming, abort
    - [Tool/Function Calling](examples/tools.md) -- Manual tools, `zodFunction`, `runTools` automation, Responses API tools
    - [Structured Output](examples/structured-output.md) -- `zodResponseFormat`, `zodTextFormat`, refusal handling
    - [Embeddings, Vision & Audio](examples/embeddings-vision-audio.md) -- Semantic search, image analysis, transcription, TTS, batch processing
    - [Quick API Reference](reference.md) -- Model IDs, method signatures, error types, streaming events
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    The official OpenAI SDK provides **direct, low-level access** to OpenAI's full API surface. It is the thinnest possible wrapper over the REST API, auto-generated from OpenAI's OpenAPI specification using Stainless.
    
    **Core principles:**
    
    1. **Direct API access** -- No abstractions or provider layers. You get the exact API that OpenAI documents, with full TypeScript types. Every API feature is available immediately when OpenAI releases it.
    2. **Two API paradigms** -- The **Responses API** (`client.responses.create()`) is the newer, recommended API with built-in tools and server-side state. The **Chat Completions API** (`client.chat.completions.create()`) remains fully supported for stateless chat flows.
    3. **Built-in resilience** -- The SDK handles retries (2 by default on 429/5xx), timeouts (10 min default), and auto-pagination out of the box.
    4. **Streaming as a first-class pattern** -- Use `stream: true` for SSE-based streaming, `.stream()` helper for event-based consumption, or `for await...of` for simple iteration.
    5. **Type-safe structured outputs** -- `zodResponseFormat()` and `client.chat.completions.parse()` convert Zod schemas to JSON Schema and parse responses, giving you validated, typed objects.
    
    **When to use the OpenAI SDK directly:**
    
    - You only use OpenAI models and want the simplest, most direct integration
    - You need access to OpenAI-specific features (Responses API, Batch, Realtime)
    - You want minimal dependencies and zero abstraction overhead
    - You need the latest API features on day one
    
    **When NOT to use:**
    
    - You need to switch between providers (OpenAI, Anthropic, Google) -- use a unified provider SDK
    - You want React-specific chat UI hooks -- use a framework-integrated AI SDK
    - You want a higher-level agent framework -- consider OpenAI Agents SDK (`@openai/agents`)
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Client Setup
    
    Initialize the OpenAI client. It auto-reads `OPENAI_API_KEY` from the environment.
    
    ```typescript
    // lib/openai.ts -- basic setup
    import OpenAI from "openai";
    const client = new OpenAI();
    export { client };
    ```
    
    ```typescript
    // lib/openai.ts -- production configuration
    const TIMEOUT_MS = 30_000;
    const MAX_RETRIES = 3;
    const client = new OpenAI({ timeout: TIMEOUT_MS, maxRetries: MAX_RETRIES });
    ```
    
    **Why good:** Minimal setup, env var auto-detected, named constants for production settings
    
    **See:** [examples/core.md](examples/core.md) for Azure OpenAI, per-request overrides, error handling patterns
    
    ---
    
    ### Pattern 2: Chat Completions API
    
    Stateless text generation. You manage conversation history.
    
    ```typescript
    const completion = await client.chat.completions.create({
      model: "gpt-4o",
      messages: [
        { role: "developer", content: "You are a helpful coding assistant." },
        { role: "user", content: "Explain TypeScript generics." },
      ],
    });
    console.log(completion.choices[0].message.content);
    ```
    
    **Why good:** Clear message roles, `developer` message for system instructions, direct content access
    
    ```typescript
    // BAD: No developer message, no error handling
    const res = await client.chat.completions.create({
      model: "gpt-4o",
      messages: [{ role: "user", content: "do something" }],
    });
    ```
    
    **Why bad:** No system instruction means unpredictable behavior, vague prompt
    
    **See:** [examples/chat.md](examples/chat.md) for multi-turn, token tracking, output length control
    
    ---
    
    ### Pattern 3: Responses API (Recommended for New Projects)
    
    Newer API with built-in tools, server-side state, and better performance with reasoning models.
    
    ```typescript
    const response = await client.responses.create({
      model: "gpt-4o",
      instructions: "You are a coding assistant.",
      input: "What are TypeScript generics?",
    });
    console.log(response.output_text);
    ```
    
    **Why good:** Clean separation of instructions and input, `output_text` helper, simpler than messages array
    
    ```typescript
    // BAD: Using Chat Completions parameters with Responses API
    const response = await client.responses.create({
      model: "gpt-4o",
      messages: [{ role: "user", content: "Hello" }], // WRONG: use 'input'
    });
    ```
    
    **Why bad:** Responses API uses `input` and `instructions`, not `messages`
    
    #### Built-in Tools
    
    Web search (`{ type: "web_search_preview" }`), file search (`{ type: "file_search" }`), code interpreter (`{ type: "code_interpreter" }`). Chain conversations with `previous_response_id` and `store: true`.
    
    **See:** [examples/tools.md](examples/tools.md) for Responses API function calling with tool outputs
    
    ---
    
    ### Pattern 4: Streaming
    
    Use streaming for user-facing responses.
    
    ```typescript
    // Chat Completions -- stream: true with for-await
    const stream = await client.chat.completions.create({
      model: "gpt-4o",
      messages: [{ role: "user", content: "Explain async/await." }],
      stream: true,
    });
    for await (const chunk of stream) {
      const content = chunk.choices[0]?.delta?.content;
      if (content) process.stdout.write(content);
    }
    ```
    
    ```typescript
    // Event-based with .stream() helper
    const stream = client.chat.completions.stream({
      model: "gpt-4o",
      messages: [{ role: "user", content: "Tell me a story." }],
    });
    stream.on("content", (delta) => process.stdout.write(delta));
    const finalContent = await stream.finalContent();
    ```
    
    **Why good:** Progressive output for better UX, event-based API for granular control
    
    ```typescript
    // BAD: Not consuming the stream
    const stream = await client.chat.completions.create({
      model: "gpt-4o",
      messages: [{ role: "user", content: "Hello" }],
      stream: true,
    });
    // Stream never consumed -- tokens are lost
    ```
    
    **Why bad:** Stream must be consumed via iteration or event handlers, otherwise tokens are lost
    
    **See:** [examples/streaming.md](examples/streaming.md) for Responses API streaming, abort, stream methods
    
    ---
    
    ### Pattern 5: Structured Outputs with Zod
    
    Use `zodResponseFormat()` and `.parse()` for type-safe structured responses.
    
    ```typescript
    import { zodResponseFormat } from "openai/helpers/zod";
    import { z } from "zod";
    
    const CalendarEvent = z.object({
      name: z.string(),
      date: z.string(),
      participants: z.array(z.string()),
    });
    
    const completion = await client.chat.completions.parse({
      model: "gpt-4o",
      messages: [
        { role: "developer", content: "Extract event details." },
        { role: "user", content: "Alice and Bob meet next Tuesday for lunch." },
      ],
      response_format: zodResponseFormat(CalendarEvent, "calendar_event"),
    });
    
    const event = completion.choices[0].message.parsed; // Fully typed
    ```
    
    **Why good:** Auto-converts schema, validates output, fully typed result, handles refusals
    
    **See:** [examples/structured-output.md](examples/structured-output.md) for Responses API (`zodTextFormat`), refusal handling, complex schemas
    
    ---
    
    ### Pattern 6: Function Calling / Tool Use
    
    Define functions the model can call. Use `zodFunction()` for type-safe definitions.
    
    ```typescript
    import { zodFunction } from "openai/helpers/zod";
    import { z } from "zod";
    
    const GetWeatherParams = z.object({
      location: z.string().describe("City name"),
      unit: z.enum(["celsius", "fahrenheit"]).default("celsius"),
    });
    
    const completion = await client.chat.completions.parse({
      model: "gpt-4o",
      messages: [{ role: "user", content: "Weather in Paris?" }],
      tools: [zodFunction({ name: "get_weather", parameters: GetWeatherParams })],
    });
    
    const toolCall = completion.choices[0].message.tool_calls?.[0];
    if (toolCall?.type === "function") {
      console.log(toolCall.function.parsed_arguments); // Typed from Zod
    }
    ```
    
    **Why good:** `zodFunction` provides type-safe argument parsing, `.describe()` guides the model
    
    Use `runTools()` for automated tool execution loops that handle the call-respond cycle automatically.
    
    **See:** [examples/tools.md](examples/tools.md) for `runTools`, manual tool definitions, Responses API function calling
    
    ---
    
    ### Pattern 7: Embeddings, Vision & Audio
    
    - **Embeddings:** `client.embeddings.create({ model: "text-embedding-3-small", input: [...] })` -- batch multiple inputs in one call
    - **Vision:** Multi-part content array with `{ type: "image_url", image_url: { url } }` for URL or base64 images
    - **Audio:** `client.audio.transcriptions.create()` for speech-to-text, `client.audio.speech.create()` for TTS
    - **Files:** `client.files.create()` with `ReadStream`, `Buffer` (via `toFile`), or `fetch()` Response
    - **Batch API:** Upload JSONL, create batch with `client.batches.create()`, poll for completion at 50% cost
    
    **See:** [examples/embeddings-vision-audio.md](examples/embeddings-vision-audio.md) for full examples with cosine similarity, base64 images, timestamps, TTS voice instructions, batch processing
    
    ---
    
    ### Pattern 8: Error Handling
    
    Always catch `OpenAI.APIError` and its subclasses. Re-throw unexpected errors.
    
    ```typescript
    try {
      const completion = await client.chat.completions.create({
        model: "gpt-4o",
        messages: [{ role: "user", content: "Hello" }],
      });
    } catch (error) {
      if (error instanceof OpenAI.APIError) {
        console.error(
          `API Error [${error.status}]: ${error.message} (${error.request_id})`,
        );
        // Check subclasses: RateLimitError, AuthenticationError, BadRequestError, etc.
      } else {
        throw error; // Re-throw non-API errors
      }
    }
    ```
    
    **Why good:** Specific error types with status codes, request ID for debugging, re-throws unexpected errors
    
    **See:** [examples/core.md](examples/core.md) for full production error handling, stream errors, error type hierarchy
    
    </patterns>
    
    ---
    
    <performance>
    
    ## Performance Optimization
    
    ### Model Selection for Cost/Speed
    
    ```
    General purpose             -> gpt-5.4 (most capable) or gpt-4o (proven, lower cost)
    Cost-sensitive / high-vol   -> gpt-5-mini or gpt-5-nano (cheapest)
    Complex reasoning           -> gpt-5.4 or o4-mini
    Structured output           -> gpt-5.4 or gpt-4o (best schema adherence)
    Embeddings                  -> text-embedding-3-small (cheapest) or text-embedding-3-large (highest quality)
    Transcription               -> whisper-1 or gpt-4o-transcribe (higher accuracy)
    TTS                         -> tts-1 (fast) or tts-1-hd (quality) or gpt-4o-mini-tts (voice control)
    Batch processing            -> gpt-5-mini at 50% batch discount
    ```
    
    ### Key Optimization Patterns
    
    - **Track token usage** via `completion.usage` for cost visibility
    - **Check `finish_reason === "length"`** to detect truncated output
    - **Use `temperature: 0`** for deterministic output (enables caching)
    - **Use `AbortController`** to cancel long-running requests
    - **Use Batch API** for high-volume jobs at 50% cost reduction
    
    </performance>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which API to Use
    
    ```
    Building a new application?
    +-- YES -> Need built-in tools (web search, file search, code interpreter)?
    |   +-- YES -> Use Responses API (client.responses.create())
    |   +-- NO -> Need server-side conversation state?
    |       +-- YES -> Use Responses API with store: true
    |       +-- NO -> Either API works, prefer Responses for new code
    +-- Existing Chat Completions code?
        +-- Working fine? -> Keep using Chat Completions (fully supported)
        +-- Need new features? -> Consider migrating to Responses API
    ```
    
    ### Which Model to Choose
    
    ```
    What is your task?
    +-- General text generation -> gpt-5.4 (most capable) or gpt-4o (lower cost)
    +-- Fast + cheap simple tasks -> gpt-5-mini or gpt-5-nano
    +-- Complex reasoning / math -> gpt-5.4 or o4-mini
    +-- Structured output -> gpt-5.4 or gpt-4o (best schema adherence)
    +-- Vision (images) -> gpt-5.4 or gpt-4o
    +-- Embeddings -> text-embedding-3-small (default) or text-embedding-3-large
    +-- Transcription -> whisper-1 or gpt-4o-transcribe
    +-- Text-to-speech -> tts-1 (fast) or gpt-4o-mini-tts (voice instructions)
    +-- Batch processing -> gpt-5-mini (cheapest at 50% batch discount)
    ```
    
    ### Streaming vs Non-Streaming
    
    ```
    Is the response user-facing?
    +-- YES -> Use streaming (stream: true or .stream())
    |   +-- Need event-level control? -> .stream() with event handlers
    |   +-- Simple text output? -> stream: true with for await
    +-- NO -> Use non-streaming
        +-- Background processing -> client.chat.completions.create()
        +-- Structured output -> client.chat.completions.parse()
        +-- High volume -> Batch API
    ```
    
    ### When to Use This SDK vs a Provider-Agnostic SDK
    
    ```
    Do you need multiple LLM providers (OpenAI + others)?
    +-- YES -> Not this skill's scope -- use a unified provider SDK
    +-- NO -> Do you need OpenAI-specific features?
        +-- YES -> Use OpenAI SDK directly
        |   Examples: Responses API, Batch API,
        |   Realtime API, built-in web search/file search
        +-- NO -> OpenAI SDK is simplest for OpenAI-only use
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Hardcoding API keys instead of using environment variables (security breach risk)
    - Using bare `catch` blocks without checking `OpenAI.APIError` (hides API errors)
    - Not consuming streams returned by `stream: true` (tokens are silently lost)
    - Using `JSON.parse()` on completion content without `zodResponseFormat` (fragile, no validation)
    - Sending full conversation history every request when Responses API's `previous_response_id` could manage state
    
    **Medium Priority Issues:**
    
    - Not setting `maxRetries` / `timeout` for production deployments (10 min default timeout may be too long)
    - Missing `developer` role message (no system instruction = unpredictable output style)
    - Using deprecated `system` role instead of `developer` role in Chat Completions
    - Not checking `finish_reason` for `'length'` truncation
    - Ignoring `usage` data (no cost visibility)
    
    **Common Mistakes:**
    
    - Confusing Responses API (`client.responses.create()`) with Chat Completions (`client.chat.completions.create()`) parameters -- they use different shapes
    - Using `messages` parameter with Responses API (it uses `input` and `instructions`)
    - Using `response_format` with models that don't support structured outputs (need gpt-4o or later)
    - Using `max_tokens` with reasoning models (o4-mini, gpt-5.x) -- use `max_completion_tokens` instead
    - Not handling the case where `completion.choices[0].message.tool_calls` is undefined
    - Forgetting that `runTools()` defaults to max 10 completions -- set `maxChatCompletions` explicitly
    
    **Gotchas & Edge Cases:**
    
    - The SDK auto-retries on 429 (rate limit) and 5xx errors -- 2 retries by default. Disable with `maxRetries: 0` if you handle retries yourself.
    - `stream: true` returns raw SSE chunks. Use `.stream()` helper for a nicer event-based API.
    - `client.chat.completions.parse()` throws `LengthFinishReasonError` if `finish_reason` is `'length'` and `ContentFilterFinishReasonError` if `'content_filter'`.
    - Embedding responses return `Array<number>` (the SDK requests base64 by default and decodes via Float32 internally for performance). No conversion needed -- you get a plain number array.
    - File uploads support `ReadStream`, `File`, `fetch()` Response, and `toFile()` helper -- use whichever matches your data source.
    - The Responses API's `store: true` enables server-side state but also means OpenAI stores your conversations. Set `store: false` for sensitive data.
    - `developer` role replaces `system` role in newer models (gpt-4o and later).
    - Batch API has a 24h completion window and 50,000 request limit per batch.
    - Audio transcription has a 25 MB file size limit.
    - Zod schemas with `zodResponseFormat` must use `additionalProperties: false` -- the SDK handles this automatically.
    - `zodTextFormat` and `zodResponseFormat` are NOT compatible with Zod v4 -- use Zod v3.x until the SDK adds v4 support.
    - The Assistants API is deprecated (sunset August 2026) -- use the Responses API for new code.
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST use the Responses API (`client.responses.create()`) for new projects -- it provides better performance, built-in tools, and server-side conversation state)**
    
    **(You MUST use `zodResponseFormat()` from `openai/helpers/zod` for structured outputs -- do NOT manually construct JSON schemas)**
    
    **(You MUST handle errors using `OpenAI.APIError` and its subclasses -- never use bare catch blocks without error type checking)**
    
    **(You MUST configure appropriate retries and timeouts for production use -- the SDK retries 2 times by default on 429/5xx errors)**
    
    **(You MUST never hardcode API keys -- always use environment variables via `process.env.OPENAI_API_KEY`)**
    
    **Failure to follow these rules will produce insecure, unreliable, or poorly-typed AI integrations.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related