Claude Skill

ai-infrastructure-together-ai

Together AI SDK patterns for TypeScript — client setup, chat completions, streaming, structured output, function calling, embeddings, image generation, fine-tuning, and OpenAI-compatible endpoints

LLM Mart · 0 points · 0 views 0 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_ai-infrastructure-together-ai_skills_ai-infrastructure-together-ai-3a51ef5.zip · 20 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-infrastructure-together-ai/skills/ai-infrastructure-together-ai
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Together AI SDK Patterns

Quick Guide: Use the together-ai npm package to access 200+ open-source models (Llama, Qwen, Mistral, DeepSeek) via Together AI's fast inference API. The SDK mirrors the OpenAI API shape -- client.chat.completions.create() for chat, client.images.generate() for images, client.embeddings.create() for embeddings. Use response_format: { type: "json_schema" } with Zod-generated schemas for structured output. Function calling uses the same tools parameter shape as OpenAI. You can also use the OpenAI SDK directly by pointing baseURL to https://api.together.xyz/v1.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST use the together-ai package (import Together from "together-ai") -- NOT the OpenAI SDK -- unless explicitly building an OpenAI-compatible integration)

(You MUST include the JSON schema in BOTH the response_format parameter AND the system prompt when using structured output -- the model needs both)

(You MUST handle errors using Together.APIError and its subclasses -- never use bare catch blocks without error type checking)

(You MUST never hardcode API keys -- always use environment variables via process.env.TOGETHER_API_KEY)

</critical_requirements>


Auto-detection: Together AI, together-ai, together.ai, TOGETHER_API_KEY, client.chat.completions (together), client.images.generate, client.embeddings.create (together), Llama-3, Qwen3, Mistral, DeepSeek, FLUX, together.images, together.chat, together.embeddings, together.fineTuning, api.together.xyz

When to use:

  • Running open-source LLMs (Llama, Qwen, Mistral, DeepSeek) via serverless inference
  • Generating images with FLUX or Stable Diffusion models
  • Creating embeddings for RAG pipelines with open-source embedding models
  • Using function calling / tool use with open-source models
  • Extracting structured JSON output from LLM responses
  • Fine-tuning open-source models on custom data
  • Migrating from OpenAI to open-source models with minimal code changes

Key patterns covered:

  • Client initialization and configuration (retries, timeouts, logging)
  • Chat completions with open-source models (Llama, Qwen, Mistral, DeepSeek)
  • Streaming with stream: true and for await...of
  • Structured output with response_format: { type: "json_schema" } and Zod
  • Function calling / tool use with tools parameter
  • Image generation with FLUX and Stable Diffusion models
  • Embeddings API with open-source embedding models
  • Fine-tuning API (file upload, job creation, monitoring)
  • OpenAI SDK compatibility (base URL swap)
  • Error handling, retries, timeouts

When NOT to use:

  • You need OpenAI-specific features (Responses API, Batch API, Realtime API) -- use the OpenAI SDK directly
  • You want framework-specific chat UI hooks -- use a framework-integrated AI SDK
  • You only use OpenAI models and never plan to use open-source models

Examples Index





<decision_framework>

Decision Framework

Which Model to Choose

What is your task?
+-- General chat / instruction following -> Llama 3.3 70B Turbo (fast, cheap)
+-- Most capable reasoning -> DeepSeek V3.1, Qwen3.5 397B
+-- Complex math / chain-of-thought -> DeepSeek R1
+-- Function calling / tool use -> Llama 3.3 70B, Qwen3.5 9B
+-- Structured JSON output -> Qwen3.5 9B (best JSON mode support)
+-- Vision / image understanding -> Qwen3-VL-8B-Instruct
+-- Code generation -> DeepSeek V3, Qwen Coder
+-- Embeddings -> BAAI/bge-large-en-v1.5 (default)
+-- Image generation (fast) -> FLUX.1 schnell
+-- Image generation (quality) -> FLUX.2 pro, FLUX.1.1 pro

Together AI SDK vs OpenAI SDK

Do you ONLY use Together AI models?
+-- YES -> Use together-ai package (purpose-built, full API coverage)
+-- NO -> Do you also use OpenAI models?
    +-- YES -> Two options:
    |   +-- Separate SDKs: together-ai for Together, openai for OpenAI
    |   +-- OpenAI SDK only: Point baseURL to api.together.xyz/v1
    +-- NO -> Use a provider-agnostic SDK

Streaming vs Non-Streaming

Is the response user-facing?
+-- YES -> Use streaming (stream: true)
+-- NO -> Use non-streaming
    +-- Background processing -> client.chat.completions.create()
    +-- Structured output -> Non-streaming with response_format

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Hardcoding TOGETHER_API_KEY instead of using environment variables (security breach risk)
  • Using bare catch blocks without checking Together.APIError (hides API errors)
  • Not consuming streams returned by stream: true (tokens are silently lost)
  • Using JSON.parse() on completion content without response_format (fragile, model may return non-JSON)
  • Omitting the schema from the system prompt when using response_format: { type: "json_schema" } (degrades output quality)

Medium Priority Issues:

  • Not setting maxRetries / timeout for production deployments (default timeout is 1 minute)
  • Missing system role message (no system instruction means unpredictable behavior)
  • Using a model that does not support function calling with tools parameter (will silently fail or error)
  • Not checking if tool_calls is defined before accessing arguments
  • Using width/height with FLUX schnell/Kontext models (use aspect_ratio instead)

Common Mistakes:

  • Using OpenAI model names (e.g., gpt-4o) with the Together AI SDK -- Together uses Hugging Face-style IDs like meta-llama/Llama-3.3-70B-Instruct-Turbo
  • Confusing client.images.generate() (Together) with client.images.create() (OpenAI) -- different method name
  • Forgetting to use z.toJSONSchema() (Zod v4) or zodToJsonSchema() (Zod v3) to convert schemas before passing to response_format
  • Using the developer role (OpenAI-specific) instead of system role with Together AI models
  • Passing max_completion_tokens instead of max_tokens -- Together uses max_tokens

Gotchas & Edge Cases:

  • The SDK auto-retries on 429 (rate limit), 408, 409, and 5xx errors -- 2 retries by default. Disable with maxRetries: 0.
  • Model IDs are case-sensitive and follow the org/model-name format from Hugging Face.
  • Not all models support function calling. See examples/tools.md for the current supported list, or check the official docs.
  • FLUX.1 schnell and Kontext models use aspect_ratio parameter; FLUX.1 Pro and FLUX.1.1 Pro use width/height.
  • Image generation returns URLs by default. Use response_format: "base64" for inline data.
  • The response_format: { type: "json_schema" } requires telling the model to "only answer in JSON" in the system prompt -- the schema alone is not sufficient.
  • Structured output uses z.toJSONSchema() (Zod v4) -- if using Zod v3, use zodToJsonSchema() from the zod-to-json-schema package.
  • Together AI's client.images.generate() is the method name, not client.images.create() like OpenAI.
  • Fine-tuning supports LoRA and full fine-tuning. File format is JSONL with messages array per line.
  • The OpenAI compatibility endpoint (api.together.xyz/v1) supports chat, embeddings, images, vision, function calling, and structured output -- but not fine-tuning or model management.

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST use the together-ai package (import Together from "together-ai") -- NOT the OpenAI SDK -- unless explicitly building an OpenAI-compatible integration)

(You MUST include the JSON schema in BOTH the response_format parameter AND the system prompt when using structured output -- the model needs both)

(You MUST handle errors using Together.APIError and its subclasses -- never use bare catch blocks without error type checking)

(You MUST never hardcode API keys -- always use environment variables via process.env.TOGETHER_API_KEY)

Failure to follow these rules will produce insecure, unreliable, or incorrectly structured AI integrations.

</critical_reminders>

Files (skills)
  • examples
    • chat.md 3.9 KB
      # Together AI SDK -- Chat Completions Examples
      
      > Chat Completions API patterns: basic completion, multi-turn conversations, model selection, vision, and token control. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [images.md](images.md) -- Image generation, embeddings
      
      ---
      
      ## Basic Chat Completion
      
      ```typescript
      // basic-chat.ts
      import Together from "together-ai";
      
      const client = new Together();
      
      async function chat(userMessage: string): Promise<string> {
        const completion = await client.chat.completions.create({
          model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
          messages: [
            {
              role: "system",
              content: "You are a helpful assistant. Be concise.",
            },
            { role: "user", content: userMessage },
          ],
        });
      
        const content = completion.choices[0].message.content;
        if (!content) {
          throw new Error("No content in response");
        }
      
        return content;
      }
      
      const answer = await chat("What is TypeScript in one sentence?");
      console.log(answer);
      ```
      
      ---
      
      ## Multi-Turn Conversations
      
      ```typescript
      import Together from "together-ai";
      import type { Together as TogetherTypes } from "together-ai";
      
      const client = new Together();
      
      const messages: TogetherTypes.Chat.CompletionCreateParams["messages"] = [
        { role: "system", content: "You are a TypeScript expert." },
        { role: "user", content: "What is a union type?" },
      ];
      
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages,
      });
      
      // Append assistant response for next turn
      const assistantMessage = completion.choices[0].message;
      messages.push({ role: "assistant", content: assistantMessage.content ?? "" });
      messages.push({ role: "user", content: "Give me a real-world example." });
      
      const followUp = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages,
      });
      ```
      
      ---
      
      ## Controlling Output Length and Temperature
      
      ```typescript
      const MAX_TOKENS = 500;
      
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [{ role: "user", content: "Summarize this article." }],
        max_tokens: MAX_TOKENS,
        temperature: 0, // Deterministic output
      });
      
      const finishReason = completion.choices[0].finish_reason;
      if (finishReason === "length") {
        console.warn("Output was truncated -- increase max_tokens");
      }
      ```
      
      ---
      
      ## Vision -- Image from URL
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      async function analyzeImage(
        imageUrl: string,
        question: string,
      ): Promise<string> {
        const response = await client.chat.completions.create({
          model: "Qwen/Qwen3-VL-8B-Instruct",
          messages: [
            {
              role: "user",
              content: [
                { type: "text", text: question },
                { type: "image_url", image_url: { url: imageUrl } },
              ],
            },
          ],
        });
      
        return response.choices[0].message.content ?? "";
      }
      
      const description = await analyzeImage(
        "https://example.com/photo.jpg",
        "Describe what you see in this image.",
      );
      console.log(description);
      ```
      
      ---
      
      ## Vision -- Multiple Images
      
      ```typescript
      const response = await client.chat.completions.create({
        model: "Qwen/Qwen3-VL-8B-Instruct",
        messages: [
          {
            role: "user",
            content: [
              { type: "text", text: "Compare these two images." },
              {
                type: "image_url",
                image_url: { url: "https://example.com/image1.jpg" },
              },
              {
                type: "image_url",
                image_url: { url: "https://example.com/image2.jpg" },
              },
            ],
          },
        ],
      });
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • core.md 5.7 KB
      # Together AI SDK -- Setup & Configuration Examples
      
      > Client initialization, environment config, production settings, error handling, and OpenAI compatibility. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [images.md](images.md) -- Image generation, embeddings
      
      ---
      
      ## Basic Client Setup
      
      ```typescript
      // lib/together.ts
      import Together from "together-ai";
      
      // Reads TOGETHER_API_KEY from env automatically
      const client = new Together();
      
      export { client };
      ```
      
      ---
      
      ## Production Configuration
      
      ```typescript
      // lib/together.ts
      import Together from "together-ai";
      
      const TIMEOUT_MS = 30_000;
      const MAX_RETRIES = 3;
      
      const client = new Together({
        apiKey: process.env.TOGETHER_API_KEY,
        timeout: TIMEOUT_MS,
        maxRetries: MAX_RETRIES,
      });
      
      export { client };
      ```
      
      ---
      
      ## Production Error Handling
      
      ```typescript
      // error-handling.ts
      import Together from "together-ai";
      
      const TIMEOUT_MS = 30_000;
      const MAX_RETRIES = 3;
      
      const client = new Together({
        timeout: TIMEOUT_MS,
        maxRetries: MAX_RETRIES,
      });
      
      async function safeCompletion(prompt: string): Promise<string | null> {
        try {
          const completion = await client.chat.completions.create({
            model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
            messages: [
              { role: "system", content: "You are a helpful assistant." },
              { role: "user", content: prompt },
            ],
          });
      
          const content = completion.choices[0].message.content;
          if (!content) {
            throw new Error("No content in response");
          }
      
          return content;
        } catch (error) {
          if (error instanceof Together.APIError) {
            console.error(`Together API Error [${error.status}]: ${error.message}`);
      
            if (error instanceof Together.RateLimitError) {
              console.error("Rate limited. SDK will auto-retry.");
              // If we get here, all retries were exhausted
              return null;
            }
      
            if (error instanceof Together.AuthenticationError) {
              throw new Error(
                "Invalid API key. Check TOGETHER_API_KEY environment variable.",
              );
            }
      
            if (error instanceof Together.BadRequestError) {
              console.error("Invalid request parameters:", error.message);
              return null;
            }
      
            if (error instanceof Together.InternalServerError) {
              console.error("Together server error after all retries");
              return null;
            }
          }
      
          // Network/connection errors
          if (error instanceof Together.APIConnectionError) {
            console.error("Network error:", error.message);
            return null;
          }
      
          // Unknown errors should be re-thrown
          throw error;
        }
      }
      
      const result = await safeCompletion("Hello!");
      if (result) {
        console.log(result);
      } else {
        console.error("Failed to get completion");
      }
      ```
      
      ---
      
      ## Error Type Hierarchy
      
      ```typescript
      // Error class hierarchy:
      // Together.APIError (base)
      //   +-- Together.BadRequestError          (400)
      //   +-- Together.AuthenticationError      (401)
      //   +-- Together.PermissionDeniedError    (403)
      //   +-- Together.NotFoundError            (404)
      //   +-- Together.UnprocessableEntityError (422)
      //   +-- Together.RateLimitError           (429)
      //   +-- Together.InternalServerError      (>=500)
      //   +-- Together.APIConnectionError       (network)
      ```
      
      ---
      
      ## OpenAI SDK Compatibility
      
      Use the OpenAI SDK with Together AI by swapping the base URL.
      
      ```typescript
      // lib/together-openai.ts
      import OpenAI from "openai";
      
      const client = new OpenAI({
        apiKey: process.env.TOGETHER_API_KEY,
        baseURL: "https://api.together.xyz/v1",
      });
      
      // Use exactly like OpenAI SDK, but with Together model IDs
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [
          { role: "system", content: "You are a helpful assistant." },
          { role: "user", content: "Hello!" },
        ],
      });
      
      console.log(completion.choices[0].message.content);
      export { client };
      ```
      
      ---
      
      ## Per-Request Overrides
      
      ```typescript
      // Override retries and timeout for a single request
      await client.chat.completions.create(
        {
          model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
          messages: [{ role: "user", content: "Hello" }],
        },
        {
          maxRetries: 5,
          timeout: 60_000,
          signal: abortController.signal,
          headers: { "X-Custom-Header": "value" },
        },
      );
      ```
      
      ---
      
      ## Request Cancellation with AbortController
      
      ```typescript
      const controller = new AbortController();
      const ABORT_TIMEOUT_MS = 5_000;
      
      // Cancel after timeout
      setTimeout(() => controller.abort(), ABORT_TIMEOUT_MS);
      
      try {
        const completion = await client.chat.completions.create(
          {
            model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
            messages: [{ role: "user", content: "Hello" }],
          },
          { signal: controller.signal },
        );
      } catch (error) {
        if (error instanceof Error && error.name === "AbortError") {
          console.log("Request was cancelled");
        }
      }
      ```
      
      ---
      
      ## Raw Response Access
      
      ```typescript
      // Get underlying Response object
      const response = await client.chat.completions
        .create({
          model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
          messages: [{ role: "user", content: "Hello" }],
        })
        .asResponse();
      
      console.log(response.status);
      console.log(response.headers.get("x-request-id"));
      
      // Or get both data and response
      const { data, response: raw } = await client.chat.completions
        .create({
          model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
          messages: [{ role: "user", content: "Hello" }],
        })
        .withResponse();
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • images.md 7.4 KB
      # Together AI SDK -- Images & Embeddings Examples
      
      > Image generation with FLUX/Stable Diffusion and embeddings for semantic search. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      
      ---
      
      ## Basic Image Generation (FLUX schnell)
      
      ```typescript
      // image-generation.ts
      import Together from "together-ai";
      
      const client = new Together();
      
      const response = await client.images.generate({
        model: "black-forest-labs/FLUX.1-schnell",
        prompt: "A serene mountain landscape at sunset with a lake reflection",
        steps: 4,
      });
      
      console.log(response.data[0].url);
      ```
      
      ---
      
      ## High Quality Image (FLUX Pro)
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const IMAGE_WIDTH = 1024;
      const IMAGE_HEIGHT = 768;
      const IMAGE_STEPS = 25;
      
      const response = await client.images.generate({
        model: "black-forest-labs/FLUX.1.1-pro",
        prompt: "A photorealistic portrait of a robot in a garden",
        width: IMAGE_WIDTH,
        height: IMAGE_HEIGHT,
        steps: IMAGE_STEPS,
      });
      
      console.log(response.data[0].url);
      ```
      
      ---
      
      ## Multiple Image Variations
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const NUM_VARIATIONS = 4;
      
      const response = await client.images.generate({
        model: "black-forest-labs/FLUX.1-schnell",
        prompt: "A cute robot assistant helping in a modern office",
        n: NUM_VARIATIONS,
        steps: 4,
      });
      
      response.data.forEach((image, index) => {
        console.log(`Variation ${index + 1}: ${image.url}`);
      });
      ```
      
      ---
      
      ## Base64 Response Format
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const response = await client.images.generate({
        model: "black-forest-labs/FLUX.1-schnell",
        prompt: "A cat in outer space",
        response_format: "base64",
      });
      
      const base64Data = response.data[0].b64_json;
      // Use directly in <img src="data:image/png;base64,${base64Data}" />
      ```
      
      ---
      
      ## Image Editing with Reference Images (FLUX.2)
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const response = await client.images.generate({
        model: "black-forest-labs/FLUX.2-pro",
        prompt: "Replace the color of the car to blue",
        width: 1024,
        height: 768,
        reference_images: ["https://example.com/original-car.jpg"],
      });
      
      console.log(response.data[0].url);
      ```
      
      ---
      
      ## Kontext Image Editing
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const response = await client.images.generate({
        model: "black-forest-labs/FLUX.1-kontext-pro",
        prompt: "Add a party hat to the dog",
        image_url: "https://example.com/dog.jpg",
      });
      
      console.log(response.data[0].url);
      ```
      
      ---
      
      ## Image Model Selection Guide
      
      ```
      Fast generation (4 steps)      -> FLUX.1 schnell (aspect_ratio param)
      High quality                   -> FLUX.1.1 pro (width/height params)
      Latest + reference images      -> FLUX.2 pro (width/height + reference_images)
      Image editing (single ref)     -> FLUX.1 Kontext pro (aspect_ratio + image_url)
      Negative prompts               -> Stable Diffusion models (negative_prompt param)
      ```
      
      ---
      
      ## Basic Embeddings
      
      ```typescript
      // embeddings.ts
      import Together from "together-ai";
      
      const client = new Together();
      const EMBEDDING_MODEL = "BAAI/bge-large-en-v1.5";
      
      const response = await client.embeddings.create({
        model: EMBEDDING_MODEL,
        input: "TypeScript provides static type checking for JavaScript.",
      });
      
      console.log("Embedding dimensions:", response.data[0].embedding.length);
      console.log("First 5 values:", response.data[0].embedding.slice(0, 5));
      ```
      
      ---
      
      ## Batch Embeddings
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      const EMBEDDING_MODEL = "BAAI/bge-large-en-v1.5";
      
      const documents = [
        "TypeScript provides static type checking.",
        "React is a library for building user interfaces.",
        "Node.js is a JavaScript runtime built on V8.",
        "PostgreSQL is a powerful relational database.",
      ];
      
      // Batch all inputs in one call for efficiency
      const response = await client.embeddings.create({
        model: EMBEDDING_MODEL,
        input: documents,
      });
      
      const embeddings = response.data.map((item) => ({
        index: item.index,
        embedding: item.embedding,
      }));
      
      console.log(`Generated ${embeddings.length} embeddings`);
      ```
      
      ---
      
      ## Semantic Search with Cosine Similarity
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      const EMBEDDING_MODEL = "BAAI/bge-large-en-v1.5";
      const SIMILARITY_THRESHOLD = 0.7;
      const TOP_K = 3;
      
      function cosineSimilarity(a: number[], b: number[]): number {
        let dot = 0;
        let normA = 0;
        let normB = 0;
        for (let i = 0; i < a.length; i++) {
          dot += a[i] * b[i];
          normA += a[i] * a[i];
          normB += b[i] * b[i];
        }
        return dot / (Math.sqrt(normA) * Math.sqrt(normB));
      }
      
      // Index documents
      const documents = [
        "TypeScript provides static type checking for JavaScript.",
        "React is a library for building user interfaces.",
        "PostgreSQL is a powerful relational database.",
        "Docker containers package applications with dependencies.",
      ];
      
      const docEmbeddings = await client.embeddings.create({
        model: EMBEDDING_MODEL,
        input: documents,
      });
      
      const indexedDocs = documents.map((text, i) => ({
        text,
        embedding: docEmbeddings.data[i].embedding,
      }));
      
      // Search
      async function search(
        query: string,
      ): Promise<Array<{ text: string; score: number }>> {
        const queryEmbedding = await client.embeddings.create({
          model: EMBEDDING_MODEL,
          input: query,
        });
      
        const queryVector = queryEmbedding.data[0].embedding;
      
        return indexedDocs
          .map((doc) => ({
            text: doc.text,
            score: cosineSimilarity(queryVector, doc.embedding),
          }))
          .filter((r) => r.score > SIMILARITY_THRESHOLD)
          .sort((a, b) => b.score - a.score)
          .slice(0, TOP_K);
      }
      
      const results = await search("What is TypeScript?");
      results.forEach((r) => {
        console.log(`[${r.score.toFixed(3)}] ${r.text}`);
      });
      ```
      
      ---
      
      ## Fine-Tuning (File Upload + Job Creation)
      
      ```typescript
      import Together from "together-ai";
      import { createReadStream } from "node:fs";
      
      const client = new Together();
      const EPOCHS = 3;
      
      // Upload training data (JSONL format)
      const file = await client.files.upload({
        file: createReadStream("training-data.jsonl"),
        purpose: "fine-tune",
      });
      console.log(`Uploaded file: ${file.id}`);
      
      // Create fine-tuning job
      const job = await client.fineTuning.create({
        training_file: file.id,
        model: "meta-llama/Meta-Llama-3-8B-Instruct",
        n_epochs: EPOCHS,
      });
      console.log(`Fine-tuning job: ${job.id}`);
      
      // Monitor job status
      const status = await client.fineTuning.retrieve(job.id);
      console.log(`Status: ${status.status}`);
      
      // List events
      const events = await client.fineTuning.listEvents(job.id);
      console.log("Events:", events);
      ```
      
      ### Training Data Format (JSONL)
      
      Each line is a JSON object with a `messages` array:
      
      ```jsonl
      {"messages": [{"role": "system", "content": "You are helpful."}, {"role": "user", "content": "Hi"}, {"role": "assistant", "content": "Hello!"}]}
      {"messages": [{"role": "user", "content": "What is TypeScript?"}, {"role": "assistant", "content": "TypeScript is a typed superset of JavaScript."}]}
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • streaming.md 3.6 KB
      # Together AI SDK -- Streaming Examples
      
      > Streaming patterns: `stream: true` with async iterators, stream cancellation, and controller access. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat Completions API
      - [tools.md](tools.md) -- Tool/function calling
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [images.md](images.md) -- Image generation, embeddings
      
      ---
      
      ## Basic Streaming with `for await`
      
      ```typescript
      // streaming-chat.ts
      import Together from "together-ai";
      
      const client = new Together();
      
      const stream = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [
          { role: "system", content: "You are a helpful assistant." },
          { role: "user", content: "Explain async/await in TypeScript." },
        ],
        stream: true,
      });
      
      for await (const chunk of stream) {
        const content = chunk.choices[0]?.delta?.content;
        if (content) {
          process.stdout.write(content);
        }
      }
      console.log(); // newline
      ```
      
      ---
      
      ## Collecting Full Response from Stream
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      async function streamToString(prompt: string): Promise<string> {
        const stream = await client.chat.completions.create({
          model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
          messages: [
            { role: "system", content: "You are a helpful assistant." },
            { role: "user", content: prompt },
          ],
          stream: true,
        });
      
        const parts: string[] = [];
      
        for await (const chunk of stream) {
          const content = chunk.choices[0]?.delta?.content;
          if (content) {
            parts.push(content);
            process.stdout.write(content); // Show progress
          }
        }
        console.log(); // newline
      
        return parts.join("");
      }
      
      const result = await streamToString("Explain promises in JavaScript.");
      console.log("Total length:", result.length);
      ```
      
      ---
      
      ## Stream Cancellation
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const stream = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [{ role: "user", content: "Tell me a long story." }],
        stream: true,
      });
      
      let tokenCount = 0;
      const MAX_TOKENS_TO_READ = 100;
      
      for await (const chunk of stream) {
        const content = chunk.choices[0]?.delta?.content;
        if (content) {
          process.stdout.write(content);
          tokenCount++;
          if (tokenCount >= MAX_TOKENS_TO_READ) {
            // Cancel the stream by using the controller
            stream.controller.abort();
            break;
          }
        }
      }
      console.log("\nStream cancelled after", tokenCount, "chunks");
      ```
      
      ---
      
      ## Streaming with Tool Calls
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const stream = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [{ role: "user", content: "What is the weather in Tokyo?" }],
        tools: [
          {
            type: "function",
            function: {
              name: "get_weather",
              description: "Get weather for a city",
              parameters: {
                type: "object",
                properties: {
                  location: { type: "string" },
                },
                required: ["location"],
              },
            },
          },
        ],
        stream: true,
      });
      
      for await (const chunk of stream) {
        const toolCalls = chunk.choices[0]?.delta?.tool_calls ?? [];
        for (const toolCall of toolCalls) {
          console.log("Tool call delta:", toolCall);
        }
      }
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • structured-output.md 5.4 KB
      # Together AI SDK -- Structured Output Examples
      
      > Structured output patterns: JSON schema mode with Zod, regex mode, vision with JSON, and complex schemas. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [tools.md](tools.md) -- Tool/function calling
      - [images.md](images.md) -- Image generation, embeddings
      
      ---
      
      ## JSON Schema Mode with Zod
      
      ```typescript
      // structured-output.ts
      import Together from "together-ai";
      import { z } from "zod";
      
      const client = new Together();
      
      const VoiceNoteSchema = z.object({
        title: z.string().describe("A title for the voice note"),
        summary: z.string().describe("A one sentence summary"),
        actionItems: z
          .array(z.string())
          .describe("A list of action items from the note"),
      });
      
      type VoiceNote = z.infer<typeof VoiceNoteSchema>;
      
      const jsonSchema = z.toJSONSchema(VoiceNoteSchema);
      
      async function extractVoiceNote(transcript: string): Promise<VoiceNote> {
        const completion = await client.chat.completions.create({
          model: "Qwen/Qwen3.5-9B",
          messages: [
            {
              role: "system",
              content: `Extract structured data from the transcript. Only answer in JSON. Follow this schema: ${JSON.stringify(jsonSchema)}`,
            },
            { role: "user", content: transcript },
          ],
          response_format: {
            type: "json_schema",
            json_schema: {
              name: "voice_note",
              schema: jsonSchema,
            },
          },
        });
      
        return JSON.parse(completion.choices[0].message.content ?? "{}");
      }
      
      const note = await extractVoiceNote(
        "Need to buy groceries today and schedule a meeting with the team for Friday.",
      );
      console.log(note.title);
      console.log(note.actionItems);
      ```
      
      ---
      
      ## Complex Nested Schema
      
      ```typescript
      import Together from "together-ai";
      import { z } from "zod";
      
      const client = new Together();
      
      const ArticleSummary = z.object({
        title: z.string(),
        summary: z.string(),
        keyPoints: z.array(z.string()),
        sentiment: z.enum(["positive", "negative", "neutral"]),
        topics: z.array(
          z.object({
            name: z.string(),
            relevance: z.number().describe("Relevance score 0-1"),
          }),
        ),
      });
      
      const jsonSchema = z.toJSONSchema(ArticleSummary);
      
      const completion = await client.chat.completions.create({
        model: "Qwen/Qwen3.5-9B",
        messages: [
          {
            role: "system",
            content: `Extract article summary. Only answer in JSON. Schema: ${JSON.stringify(jsonSchema)}`,
          },
          {
            role: "user",
            content:
              "TypeScript 5.8 brings improved type inference and better error messages, making everyday coding more productive.",
          },
        ],
        response_format: {
          type: "json_schema",
          json_schema: { name: "article_summary", schema: jsonSchema },
        },
      });
      
      const article = JSON.parse(completion.choices[0].message.content ?? "{}");
      console.log(`Title: ${article.title}`);
      console.log(`Sentiment: ${article.sentiment}`);
      ```
      
      ---
      
      ## Regex Mode (Constrained Output)
      
      Constrain output to a regex pattern for classification tasks.
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const MAX_TOKENS = 10;
      
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        temperature: 0.2,
        max_tokens: MAX_TOKENS,
        messages: [
          {
            role: "system",
            content:
              "Classify the sentiment of the text as positive, neutral, or negative.",
          },
          { role: "user", content: "Wow. I loved the movie!" },
        ],
        response_format: {
          type: "regex",
          // @ts-ignore -- regex type not in SDK types yet
          pattern: "(positive|neutral|negative)",
        },
      });
      
      console.log(completion.choices[0].message.content);
      // Output: "positive"
      ```
      
      ---
      
      ## Vision Model with JSON Output
      
      Extract structured data from images.
      
      ```typescript
      import Together from "together-ai";
      import { z } from "zod";
      
      const client = new Together();
      
      const ImageDescription = z.object({
        description: z.string().describe("What the image shows"),
        objectCount: z.number().describe("Number of main objects"),
        dominantColors: z.array(z.string()).describe("Main colors in the image"),
      });
      
      const jsonSchema = z.toJSONSchema(ImageDescription);
      
      const completion = await client.chat.completions.create({
        model: "Qwen/Qwen3-VL-8B-Instruct",
        messages: [
          {
            role: "user",
            content: [
              {
                type: "text",
                text: `Describe this image. Only answer in JSON. Schema: ${JSON.stringify(jsonSchema)}`,
              },
              {
                type: "image_url",
                image_url: { url: "https://example.com/photo.jpg" },
              },
            ],
          },
        ],
        response_format: {
          type: "json_schema",
          json_schema: { name: "image_description", schema: jsonSchema },
        },
      });
      
      const description = JSON.parse(completion.choices[0].message.content ?? "{}");
      console.log(description);
      ```
      
      ---
      
      ## Zod v3 vs Zod v4
      
      ```typescript
      // Zod v4 (recommended): Use built-in z.toJSONSchema()
      import { z } from "zod";
      const schema = z.object({ name: z.string() });
      const jsonSchema = z.toJSONSchema(schema); // Built-in
      
      // Zod v3: Use zodToJsonSchema from separate package
      import { z } from "zod";
      import { zodToJsonSchema } from "zod-to-json-schema";
      const schema = z.object({ name: z.string() });
      const jsonSchema = zodToJsonSchema(schema); // External package
      ```
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
    • tools.md 7.6 KB
      # Together AI SDK -- Tool/Function Calling Examples
      
      > Function calling patterns: tool definitions, tool_choice control, multi-step tool loops, parallel calls. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, error handling
      - [chat.md](chat.md) -- Chat Completions API
      - [streaming.md](streaming.md) -- Streaming responses
      - [structured-output.md](structured-output.md) -- Structured outputs with Zod
      - [images.md](images.md) -- Image generation, embeddings
      
      ---
      
      ## Basic Function Calling
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [{ role: "user", content: "What is the weather in Tokyo?" }],
        tools: [
          {
            type: "function",
            function: {
              name: "get_weather",
              description: "Get the current weather for a location",
              parameters: {
                type: "object",
                properties: {
                  location: { type: "string", description: "City name" },
                  unit: {
                    type: "string",
                    enum: ["celsius", "fahrenheit"],
                    description: "Temperature unit",
                  },
                },
                required: ["location"],
                additionalProperties: false,
              },
              strict: true,
            },
          },
        ],
      });
      
      const toolCall = completion.choices[0].message.tool_calls?.[0];
      if (toolCall) {
        const args = JSON.parse(toolCall.function.arguments);
        console.log(`Call ${toolCall.function.name} with:`, args);
      }
      ```
      
      ---
      
      ## Multiple Tools
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [
          {
            role: "system",
            content: "You help users by calling available tools.",
          },
          {
            role: "user",
            content: "What is the weather in London and search for restaurants?",
          },
        ],
        tools: [
          {
            type: "function",
            function: {
              name: "get_weather",
              description: "Get current weather for a city",
              parameters: {
                type: "object",
                properties: {
                  location: { type: "string", description: "City name" },
                },
                required: ["location"],
                additionalProperties: false,
              },
            },
          },
          {
            type: "function",
            function: {
              name: "search_restaurants",
              description: "Search for restaurants in a city",
              parameters: {
                type: "object",
                properties: {
                  city: { type: "string", description: "City to search" },
                  cuisine: { type: "string", description: "Cuisine type" },
                },
                required: ["city"],
                additionalProperties: false,
              },
            },
          },
        ],
      });
      
      // Model may return one or more tool calls
      const toolCalls = completion.choices[0].message.tool_calls ?? [];
      for (const toolCall of toolCalls) {
        const args = JSON.parse(toolCall.function.arguments);
        console.log(`Call ${toolCall.function.name}:`, args);
      }
      ```
      
      ---
      
      ## Controlling Tool Invocation with `tool_choice`
      
      ```typescript
      // Force a specific function
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [{ role: "user", content: "Tell me about Paris" }],
        tools: [
          /* ... */
        ],
        tool_choice: {
          type: "function",
          function: { name: "get_weather" },
        },
      });
      
      // Force at least one tool call (any tool)
      const required = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [{ role: "user", content: "Hello" }],
        tools: [
          /* ... */
        ],
        tool_choice: "required",
      });
      
      // Disable tool calling for this request
      const noTools = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [{ role: "user", content: "Hello" }],
        tools: [
          /* ... */
        ],
        tool_choice: "none",
      });
      ```
      
      ---
      
      ## Multi-Step Tool Loop
      
      Execute tool calls and feed results back to the model.
      
      ```typescript
      import Together from "together-ai";
      
      const client = new Together();
      
      // Tool implementations
      async function getWeather(location: string): Promise<string> {
        return JSON.stringify({ location, temperature: 22, condition: "sunny" });
      }
      
      async function searchDatabase(query: string): Promise<string> {
        return JSON.stringify({ results: [{ id: 1, title: `Result: ${query}` }] });
      }
      
      const toolImplementations: Record<
        string,
        (args: Record<string, string>) => Promise<string>
      > = {
        get_weather: (args) => getWeather(args.location),
        search_database: (args) => searchDatabase(args.query),
      };
      
      const tools = [
        {
          type: "function" as const,
          function: {
            name: "get_weather",
            description: "Get weather for a city",
            parameters: {
              type: "object" as const,
              properties: {
                location: { type: "string", description: "City name" },
              },
              required: ["location"],
              additionalProperties: false,
            },
          },
        },
        {
          type: "function" as const,
          function: {
            name: "search_database",
            description: "Search the database",
            parameters: {
              type: "object" as const,
              properties: {
                query: { type: "string", description: "Search query" },
              },
              required: ["query"],
              additionalProperties: false,
            },
          },
        },
      ];
      
      // Initial request
      const messages: Array<{
        role: "system" | "user" | "assistant" | "tool";
        content: string | null;
        tool_calls?: Array<{
          id: string;
          type: "function";
          function: { name: string; arguments: string };
        }>;
        tool_call_id?: string;
        name?: string;
      }> = [
        { role: "system", content: "Help users with weather and database queries." },
        { role: "user", content: "What is the weather in London?" },
      ];
      
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages,
        tools,
      });
      
      const assistantMessage = completion.choices[0].message;
      
      // If model wants to call tools, execute them and continue
      if (assistantMessage.tool_calls && assistantMessage.tool_calls.length > 0) {
        // Add assistant message with tool calls
        messages.push({
          role: "assistant",
          content: assistantMessage.content,
          tool_calls: assistantMessage.tool_calls,
        });
      
        // Execute each tool and add results
        for (const toolCall of assistantMessage.tool_calls) {
          const args = JSON.parse(toolCall.function.arguments);
          const impl = toolImplementations[toolCall.function.name];
          const result = impl ? await impl(args) : "Unknown tool";
      
          messages.push({
            role: "tool",
            content: result,
            tool_call_id: toolCall.id,
            name: toolCall.function.name,
          });
        }
      
        // Get final response without tools
        const finalCompletion = await client.chat.completions.create({
          model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
          messages,
        });
      
        console.log("Final answer:", finalCompletion.choices[0].message.content);
      }
      ```
      
      ---
      
      ## Supported Models for Function Calling
      
      Not all Together AI models support function calling. Currently supported:
      
      - `meta-llama/Llama-3.3-70B-Instruct-Turbo`
      - `meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8`
      - `Qwen/Qwen3.5-397B-A17B`
      - `Qwen/Qwen3.5-9B`
      - `Qwen/Qwen3-Next-80B-A3B-Instruct`
      - `Qwen/Qwen2.5-7B-Instruct-Turbo`
      - `deepseek-ai/DeepSeek-V3`
      - `deepseek-ai/DeepSeek-R1`
      - `mistralai/Mistral-Small-24B-Instruct-2501`
      
      Check [Together AI docs](https://docs.together.ai/docs/function-calling) for the latest model support.
      
      ---
      
      _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
      
  • reference.md 8.4 KB
    # Together AI SDK Quick Reference
    
    > Client configuration, model IDs, API methods, error types, and image parameters. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Package Installation
    
    ```bash
    # Core package (always required)
    npm install together-ai
    
    # For structured outputs (optional but recommended)
    npm install zod
    ```
    
    ---
    
    ## Client Configuration
    
    ```typescript
    import Together from "together-ai";
    
    const client = new Together({
      apiKey: process.env.TOGETHER_API_KEY, // Auto-reads from env if not set
      timeout: 30_000, // Request timeout in ms (default: 60_000 = 1 min)
      maxRetries: 3, // Retry count on 429/5xx (default: 2)
      baseURL: "https://api.together.xyz/v1", // Override for proxies
      logLevel: "off", // 'debug' | 'info' | 'warn' | 'error' | 'off'
    });
    ```
    
    ### Environment Variables
    
    | Variable           | Purpose                 |
    | ------------------ | ----------------------- |
    | `TOGETHER_API_KEY` | API key (auto-detected) |
    
    ---
    
    ## Model IDs
    
    ### Chat / Language Models
    
    | Model ID                                            | Use Case                    | Notes                         |
    | --------------------------------------------------- | --------------------------- | ----------------------------- |
    | `meta-llama/Llama-3.3-70B-Instruct-Turbo`           | General purpose, fast       | Best balance of speed/quality |
    | `meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8` | Latest Llama 4              | MoE architecture              |
    | `Qwen/Qwen3.5-9B`                                   | JSON mode, function calling | Excellent structured output   |
    | `Qwen/Qwen3.5-397B-A17B`                            | Most capable Qwen           | MoE, hybrid reasoning         |
    | `deepseek-ai/DeepSeek-V3.1`                         | Most capable open-source    | Strong reasoning              |
    | `deepseek-ai/DeepSeek-R1`                           | Complex reasoning           | Chain-of-thought reasoning    |
    | `mistralai/Mistral-Small-24B-Instruct-2501`         | Fast, function calling      | Good tool use support         |
    | `google/gemma-3n-E4B-it`                            | Ultra-lightweight           | Cheapest option               |
    
    ### Vision / Multimodal Models
    
    | Model ID                                         | Use Case            |
    | ------------------------------------------------ | ------------------- |
    | `Qwen/Qwen3-VL-8B-Instruct`                      | Image understanding |
    | `meta-llama/Llama-3.2-11B-Vision-Instruct-Turbo` | Vision + chat       |
    
    ### Embedding Models
    
    | Model ID                                    | Use Case               |
    | ------------------------------------------- | ---------------------- |
    | `BAAI/bge-large-en-v1.5`                    | General English        |
    | `WhereIsAI/UAE-Large-V1`                    | High quality           |
    | `intfloat/multilingual-e5-large-instruct`   | Multilingual           |
    | `togethercomputer/m2-bert-80M-8k-retrieval` | Long context retrieval |
    
    ### Image Generation Models
    
    | Model ID                               | Use Case               | Size Control       |
    | -------------------------------------- | ---------------------- | ------------------ |
    | `black-forest-labs/FLUX.1-schnell`     | Fast generation        | `aspect_ratio`     |
    | `black-forest-labs/FLUX.1.1-pro`       | High quality           | `width` / `height` |
    | `black-forest-labs/FLUX.2-pro`         | Latest, reference imgs | `width` / `height` |
    | `black-forest-labs/FLUX.1-kontext-pro` | Image editing          | `aspect_ratio`     |
    
    ---
    
    ## API Methods Reference
    
    ### Chat Completions
    
    ```typescript
    const completion = await client.chat.completions.create({
      model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", // Required
      messages: [], // Required: ChatCompletionMessageParam[]
      temperature: 0.7, // 0-2 (default: 1)
      max_tokens: 1000, // Max output tokens
      top_p: 1, // Nucleus sampling
      tools: [], // Function calling tools
      tool_choice: "auto", // 'auto' | 'required' | 'none' | { type: 'function', function: { name } }
      response_format: undefined, // { type: 'json_schema', json_schema: { name, schema } }
      stream: false, // Enable streaming
      stop: undefined, // string | string[] -- stop sequences
    });
    ```
    
    ### Images
    
    ```typescript
    const response = await client.images.generate({
      model: "black-forest-labs/FLUX.1-schnell", // Required
      prompt: "", // Required (except Kling)
      width: 1024, // For FLUX Pro/1.1 Pro/Dev
      height: 1024, // For FLUX Pro/1.1 Pro/Dev
      aspect_ratio: undefined, // For FLUX schnell/Kontext: "1:1", "16:9", etc.
      steps: 4, // 1-50
      n: 1, // 1-4 variations
      seed: undefined, // Reproducibility
      negative_prompt: "", // Unwanted elements
      response_format: undefined, // "base64" for inline data
      reference_images: [], // FLUX.2, Google models
      image_url: undefined, // Kontext single reference
      disable_safety_checker: false, // NSFW filter toggle
    });
    ```
    
    ### Embeddings
    
    ```typescript
    const response = await client.embeddings.create({
      model: "BAAI/bge-large-en-v1.5", // Required
      input: "", // string or string[]
    });
    ```
    
    ### Fine-Tuning
    
    ```typescript
    // Upload training data
    const file = await client.files.upload({
      file: readStream, // ReadStream or File
      purpose: "fine-tune",
    });
    
    // Create fine-tuning job
    const job = await client.fineTuning.create({
      training_file: file.id, // Required
      model: "meta-llama/Meta-Llama-3-8B-Instruct", // Required
      n_epochs: 3,
    });
    
    // Monitor job
    const status = await client.fineTuning.retrieve(job.id);
    const events = await client.fineTuning.listEvents(job.id);
    
    // List, cancel, delete
    const jobs = await client.fineTuning.list();
    await client.fineTuning.cancel(job.id);
    await client.fineTuning.delete(job.id);
    ```
    
    ### Files
    
    ```typescript
    // Upload
    const file = await client.files.upload({
      file: readStream, // ReadStream or File
      purpose: "fine-tune", // 'fine-tune'
    });
    
    // List / Retrieve / Delete
    const files = await client.files.list();
    const fileInfo = await client.files.retrieve("file-abc123");
    await client.files.delete("file-abc123");
    ```
    
    ---
    
    ## Error Types
    
    | Error Class                 | HTTP Status | Auto-Retried? |
    | --------------------------- | ----------- | ------------- |
    | `BadRequestError`           | 400         | No            |
    | `AuthenticationError`       | 401         | No            |
    | `PermissionDeniedError`     | 403         | No            |
    | `NotFoundError`             | 404         | No            |
    | `UnprocessableEntityError`  | 422         | No            |
    | `RateLimitError`            | 429         | Yes           |
    | `InternalServerError`       | >= 500      | Yes           |
    | `APIConnectionError`        | N/A         | Yes           |
    | `APIConnectionTimeoutError` | N/A         | Yes           |
    
    All errors extend `Together.APIError` with properties:
    
    - `.status` -- HTTP status code
    - `.message` -- Error message
    
    ---
    
    ## OpenAI Compatibility
    
    ```typescript
    import OpenAI from "openai";
    
    const client = new OpenAI({
      apiKey: process.env.TOGETHER_API_KEY,
      baseURL: "https://api.together.xyz/v1",
    });
    
    // Use exactly like OpenAI SDK, with Together AI model IDs
    const completion = await client.chat.completions.create({
      model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
      messages: [{ role: "user", content: "Hello" }],
    });
    ```
    
    **Compatible endpoints:** chat completions, embeddings, images, vision, function calling, structured output
    
    **NOT compatible:** fine-tuning management, model listing, Together-specific endpoints
    
    ---
    
    ## Message Roles
    
    | Role        | Description        | Notes                         |
    | ----------- | ------------------ | ----------------------------- |
    | `system`    | System instruction | Use `system`, NOT `developer` |
    | `user`      | User input         |                               |
    | `assistant` | Model response     |                               |
    | `tool`      | Tool result        | Used in multi-step tool flows |
    
    ---
    
    ## Response Format Options
    
    | Type          | Shape                                                         | Use Case                |
    | ------------- | ------------------------------------------------------------- | ----------------------- |
    | `json_schema` | `{ type: "json_schema", json_schema: { name, schema } }`      | Structured JSON output  |
    | `json_object` | `{ type: "json_object" }`                                     | Freeform JSON           |
    | `regex`       | `{ type: "regex", pattern: "(positive\|negative\|neutral)" }` | Constrained text output |
    
  • SKILL.md 19.4 KB
    ---
    name: ai-infrastructure-together-ai
    description: Together AI SDK patterns for TypeScript — client setup, chat completions, streaming, structured output, function calling, embeddings, image generation, fine-tuning, and OpenAI-compatible endpoints
    ---
    
    # Together AI SDK Patterns
    
    > **Quick Guide:** Use the `together-ai` npm package to access 200+ open-source models (Llama, Qwen, Mistral, DeepSeek) via Together AI's fast inference API. The SDK mirrors the OpenAI API shape -- `client.chat.completions.create()` for chat, `client.images.generate()` for images, `client.embeddings.create()` for embeddings. Use `response_format: { type: "json_schema" }` with Zod-generated schemas for structured output. Function calling uses the same `tools` parameter shape as OpenAI. You can also use the OpenAI SDK directly by pointing `baseURL` to `https://api.together.xyz/v1`.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST use the `together-ai` package (`import Together from "together-ai"`) -- NOT the OpenAI SDK -- unless explicitly building an OpenAI-compatible integration)**
    
    **(You MUST include the JSON schema in BOTH the `response_format` parameter AND the system prompt when using structured output -- the model needs both)**
    
    **(You MUST handle errors using `Together.APIError` and its subclasses -- never use bare catch blocks without error type checking)**
    
    **(You MUST never hardcode API keys -- always use environment variables via `process.env.TOGETHER_API_KEY`)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** Together AI, together-ai, together.ai, TOGETHER_API_KEY, client.chat.completions (together), client.images.generate, client.embeddings.create (together), Llama-3, Qwen3, Mistral, DeepSeek, FLUX, together.images, together.chat, together.embeddings, together.fineTuning, api.together.xyz
    
    **When to use:**
    
    - Running open-source LLMs (Llama, Qwen, Mistral, DeepSeek) via serverless inference
    - Generating images with FLUX or Stable Diffusion models
    - Creating embeddings for RAG pipelines with open-source embedding models
    - Using function calling / tool use with open-source models
    - Extracting structured JSON output from LLM responses
    - Fine-tuning open-source models on custom data
    - Migrating from OpenAI to open-source models with minimal code changes
    
    **Key patterns covered:**
    
    - Client initialization and configuration (retries, timeouts, logging)
    - Chat completions with open-source models (Llama, Qwen, Mistral, DeepSeek)
    - Streaming with `stream: true` and `for await...of`
    - Structured output with `response_format: { type: "json_schema" }` and Zod
    - Function calling / tool use with `tools` parameter
    - Image generation with FLUX and Stable Diffusion models
    - Embeddings API with open-source embedding models
    - Fine-tuning API (file upload, job creation, monitoring)
    - OpenAI SDK compatibility (base URL swap)
    - Error handling, retries, timeouts
    
    **When NOT to use:**
    
    - You need OpenAI-specific features (Responses API, Batch API, Realtime API) -- use the OpenAI SDK directly
    - You want framework-specific chat UI hooks -- use a framework-integrated AI SDK
    - You only use OpenAI models and never plan to use open-source models
    
    ---
    
    ## Examples Index
    
    - [Core: Setup & Configuration](examples/core.md) -- Client init, production config, error handling, OpenAI compatibility
    - [Chat Completions](examples/chat.md) -- Basic chat, multi-turn, model selection, vision
    - [Streaming](examples/streaming.md) -- Async iteration, stream cancellation
    - [Tool/Function Calling](examples/tools.md) -- Tool definitions, multi-step tool loops
    - [Structured Output](examples/structured-output.md) -- JSON mode, Zod schemas, regex mode
    - [Images & Embeddings](examples/images.md) -- FLUX image generation, embedding models, semantic search
    - [Quick API Reference](reference.md) -- Model IDs, method signatures, error types
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    Together AI provides **fast serverless inference for open-source models**. The TypeScript SDK (`together-ai`) is auto-generated with Stainless and mirrors the OpenAI API shape, making migration straightforward.
    
    **Core principles:**
    
    1. **OpenAI-compatible API shape** -- Same `client.chat.completions.create()` pattern, same `messages` array, same `tools` parameter. Switching from OpenAI is often just changing the import and model name.
    2. **Open-source model access** -- Run Llama, Qwen, Mistral, DeepSeek, and 200+ other models without managing infrastructure. Models are identified by their Hugging Face-style IDs (e.g., `meta-llama/Llama-3.3-70B-Instruct-Turbo`).
    3. **Multi-modal support** -- Chat completions, image generation (FLUX, Stable Diffusion), embeddings, audio, and video -- all through one SDK.
    4. **Structured output via JSON Schema** -- Pass a JSON schema in `response_format` and include it in the system prompt. Use Zod's `z.toJSONSchema()` to generate schemas from TypeScript types.
    5. **Fine-tuning open-source models** -- Upload JSONL data, create LoRA or full fine-tuning jobs, and deploy custom models -- all via the API.
    
    **When to use Together AI:**
    
    - You want to use open-source models with fast serverless inference
    - You need cost-effective inference (often cheaper than proprietary APIs)
    - You want to fine-tune open-source models on your data
    - You need image generation with FLUX models
    - You want OpenAI API compatibility for easy migration
    
    **When NOT to use:**
    
    - You need OpenAI-specific features (Responses API, Batch API, Realtime) -- use the OpenAI SDK
    - You need Anthropic or Google-specific features -- use their respective SDKs
    - You want a provider-agnostic SDK -- use a unified provider framework
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Client Setup
    
    Initialize the Together client. It reads `TOGETHER_API_KEY` from the environment.
    
    ```typescript
    // lib/together.ts -- basic setup
    import Together from "together-ai";
    const client = new Together();
    export { client };
    ```
    
    ```typescript
    // lib/together.ts -- production configuration
    const TIMEOUT_MS = 30_000;
    const MAX_RETRIES = 3;
    
    const client = new Together({
      apiKey: process.env.TOGETHER_API_KEY,
      timeout: TIMEOUT_MS,
      maxRetries: MAX_RETRIES,
    });
    export { client };
    ```
    
    **Why good:** Minimal setup, env var auto-detected, named constants for production settings
    
    ```typescript
    // BAD: Hardcoded API key
    const client = new Together({
      apiKey: "sk-abc123...",
    });
    ```
    
    **Why bad:** Hardcoded keys get leaked in version control, security breach risk
    
    **See:** [examples/core.md](examples/core.md) for error handling, OpenAI compatibility, per-request overrides
    
    ---
    
    ### Pattern 2: Chat Completions
    
    Stateless text generation with open-source models.
    
    ```typescript
    const completion = await client.chat.completions.create({
      model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
      messages: [
        { role: "system", content: "You are a helpful coding assistant." },
        { role: "user", content: "Explain TypeScript generics." },
      ],
    });
    console.log(completion.choices[0].message.content);
    ```
    
    **Why good:** Clear message roles, system message for behavior control, direct content access
    
    ```typescript
    // BAD: No system message, no model specified
    const res = await client.chat.completions.create({
      messages: [{ role: "user", content: "do something" }],
    });
    ```
    
    **Why bad:** Missing `model` field will error, no system instruction means unpredictable behavior
    
    **See:** [examples/chat.md](examples/chat.md) for multi-turn, vision models, model selection guide
    
    ---
    
    ### Pattern 3: Streaming
    
    Use streaming for user-facing responses.
    
    ```typescript
    const stream = await client.chat.completions.create({
      model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
      messages: [{ role: "user", content: "Explain async/await." }],
      stream: true,
    });
    for await (const chunk of stream) {
      const content = chunk.choices[0]?.delta?.content;
      if (content) process.stdout.write(content);
    }
    ```
    
    **Why good:** Progressive output for better UX, standard async iterator pattern
    
    ```typescript
    // BAD: Not consuming the stream
    const stream = await client.chat.completions.create({
      model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
      messages: [{ role: "user", content: "Hello" }],
      stream: true,
    });
    // Stream never consumed -- tokens are lost
    ```
    
    **Why bad:** Stream must be consumed via iteration, otherwise tokens are silently lost
    
    **See:** [examples/streaming.md](examples/streaming.md) for stream cancellation, controller access
    
    ---
    
    ### Pattern 4: Structured Output with JSON Schema
    
    Use `response_format: { type: "json_schema" }` with Zod-generated schemas.
    
    ```typescript
    import Together from "together-ai";
    import { z } from "zod";
    
    const client = new Together();
    
    const EventSchema = z.object({
      name: z.string(),
      date: z.string(),
      participants: z.array(z.string()),
    });
    
    const jsonSchema = z.toJSONSchema(EventSchema);
    
    const completion = await client.chat.completions.create({
      model: "Qwen/Qwen3.5-9B",
      messages: [
        {
          role: "system",
          content: `Extract event details. Only answer in JSON. Follow this schema: ${JSON.stringify(jsonSchema)}`,
        },
        { role: "user", content: "Alice and Bob meet next Tuesday for lunch." },
      ],
      response_format: {
        type: "json_schema",
        json_schema: { name: "calendar_event", schema: jsonSchema },
      },
    });
    
    const event = JSON.parse(completion.choices[0].message.content ?? "{}");
    ```
    
    **Why good:** Zod generates schema, schema included in both system prompt and `response_format`, named schema object
    
    ```typescript
    // BAD: Schema only in response_format, not in system prompt
    const completion = await client.chat.completions.create({
      model: "Qwen/Qwen3.5-9B",
      messages: [{ role: "user", content: "Extract event details." }],
      response_format: {
        type: "json_schema",
        json_schema: { name: "event", schema: jsonSchema },
      },
    });
    ```
    
    **Why bad:** Model needs the schema in the system prompt AND `response_format` for reliable structured output -- omitting the prompt instruction degrades output quality
    
    **See:** [examples/structured-output.md](examples/structured-output.md) for regex mode, vision with JSON, complex schemas
    
    ---
    
    ### Pattern 5: Function Calling / Tool Use
    
    Define functions the model can call. Same `tools` parameter shape as OpenAI.
    
    ```typescript
    const completion = await client.chat.completions.create({
      model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
      messages: [{ role: "user", content: "Weather in Paris?" }],
      tools: [
        {
          type: "function",
          function: {
            name: "get_weather",
            description: "Get current weather for a location",
            parameters: {
              type: "object",
              properties: {
                location: { type: "string", description: "City name" },
              },
              required: ["location"],
              additionalProperties: false,
            },
            strict: true,
          },
        },
      ],
    });
    
    const toolCall = completion.choices[0].message.tool_calls?.[0];
    if (toolCall) {
      const args = JSON.parse(toolCall.function.arguments);
      console.log(`Call ${toolCall.function.name} with:`, args);
    }
    ```
    
    **Why good:** Standard OpenAI-compatible tool format, strict mode for reliable arguments, `additionalProperties: false` prevents hallucinated fields
    
    **See:** [examples/tools.md](examples/tools.md) for multi-step tool loops, `tool_choice`, parallel calls, supported models
    
    ---
    
    ### Pattern 6: Image Generation
    
    Generate images with FLUX and Stable Diffusion models.
    
    ```typescript
    const response = await client.images.generate({
      model: "black-forest-labs/FLUX.1-schnell",
      prompt: "A serene mountain landscape at sunset with a lake reflection",
      steps: 4,
    });
    console.log(response.data[0].url);
    ```
    
    **Why good:** Simple API, model-specific parameters, URL response by default
    
    **See:** [examples/images.md](examples/images.md) for FLUX variants, base64, reference images, multiple variations
    
    ---
    
    ### Pattern 7: Embeddings
    
    Create embeddings for semantic search and RAG pipelines.
    
    ```typescript
    const EMBEDDING_MODEL = "BAAI/bge-large-en-v1.5";
    
    const response = await client.embeddings.create({
      model: EMBEDDING_MODEL,
      input: "TypeScript provides static type checking.",
    });
    console.log(response.data[0].embedding);
    ```
    
    **Why good:** Named model constant, simple single-input embedding, array response
    
    **See:** [examples/images.md](examples/images.md) for batch embeddings, semantic search with cosine similarity
    
    ---
    
    ### Pattern 8: Error Handling
    
    Always catch `Together.APIError` and its subclasses.
    
    ```typescript
    try {
      const completion = await client.chat.completions.create({
        model: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
        messages: [{ role: "user", content: "Hello" }],
      });
    } catch (error) {
      if (error instanceof Together.APIError) {
        console.error(`API Error [${error.status}]: ${error.message}`);
        if (error instanceof Together.RateLimitError) {
          console.error("Rate limited -- SDK will auto-retry.");
        }
        if (error instanceof Together.AuthenticationError) {
          throw new Error("Invalid API key. Check TOGETHER_API_KEY.");
        }
      } else {
        throw error; // Re-throw non-API errors
      }
    }
    ```
    
    **Why good:** Specific error types, re-throws unexpected errors, actionable error messages
    
    **See:** [examples/core.md](examples/core.md) for full production error handling, error type hierarchy
    
    </patterns>
    
    ---
    
    <performance>
    
    ## Performance Optimization
    
    ### Model Selection for Cost/Speed
    
    ```
    Fast + cheap              -> Llama 3.3 70B Turbo, Qwen3.5 9B
    Most capable              -> DeepSeek V3.1, Qwen3.5 397B
    Complex reasoning         -> DeepSeek R1
    Function calling          -> Llama 3.3 70B, Qwen3.5 9B, DeepSeek V3
    Structured output (JSON)  -> Qwen3.5 9B, Llama 3.3 70B
    Embeddings                -> BAAI/bge-large-en-v1.5 (quality), UAE-Large-V1
    Image generation (fast)   -> FLUX.1 schnell (4 steps)
    Image generation (quality)-> FLUX.2 pro, FLUX.1.1 pro
    Vision / multimodal       -> Qwen3-VL-8B-Instruct, Llama 3.2 Vision
    ```
    
    ### Key Optimization Patterns
    
    - **Use Turbo variants** for chat models -- they are optimized for Together's infrastructure
    - **Set `temperature: 0`** for deterministic output when possible
    - **Batch embedding inputs** -- pass an array of strings to `client.embeddings.create()` instead of one at a time
    - **Use `steps: 4`** for FLUX.1 schnell images (higher steps have diminishing returns)
    - **Use streaming** for user-facing responses to reduce perceived latency
    
    </performance>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Model to Choose
    
    ```
    What is your task?
    +-- General chat / instruction following -> Llama 3.3 70B Turbo (fast, cheap)
    +-- Most capable reasoning -> DeepSeek V3.1, Qwen3.5 397B
    +-- Complex math / chain-of-thought -> DeepSeek R1
    +-- Function calling / tool use -> Llama 3.3 70B, Qwen3.5 9B
    +-- Structured JSON output -> Qwen3.5 9B (best JSON mode support)
    +-- Vision / image understanding -> Qwen3-VL-8B-Instruct
    +-- Code generation -> DeepSeek V3, Qwen Coder
    +-- Embeddings -> BAAI/bge-large-en-v1.5 (default)
    +-- Image generation (fast) -> FLUX.1 schnell
    +-- Image generation (quality) -> FLUX.2 pro, FLUX.1.1 pro
    ```
    
    ### Together AI SDK vs OpenAI SDK
    
    ```
    Do you ONLY use Together AI models?
    +-- YES -> Use together-ai package (purpose-built, full API coverage)
    +-- NO -> Do you also use OpenAI models?
        +-- YES -> Two options:
        |   +-- Separate SDKs: together-ai for Together, openai for OpenAI
        |   +-- OpenAI SDK only: Point baseURL to api.together.xyz/v1
        +-- NO -> Use a provider-agnostic SDK
    ```
    
    ### Streaming vs Non-Streaming
    
    ```
    Is the response user-facing?
    +-- YES -> Use streaming (stream: true)
    +-- NO -> Use non-streaming
        +-- Background processing -> client.chat.completions.create()
        +-- Structured output -> Non-streaming with response_format
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Hardcoding `TOGETHER_API_KEY` instead of using environment variables (security breach risk)
    - Using bare `catch` blocks without checking `Together.APIError` (hides API errors)
    - Not consuming streams returned by `stream: true` (tokens are silently lost)
    - Using `JSON.parse()` on completion content without `response_format` (fragile, model may return non-JSON)
    - Omitting the schema from the system prompt when using `response_format: { type: "json_schema" }` (degrades output quality)
    
    **Medium Priority Issues:**
    
    - Not setting `maxRetries` / `timeout` for production deployments (default timeout is 1 minute)
    - Missing `system` role message (no system instruction means unpredictable behavior)
    - Using a model that does not support function calling with `tools` parameter (will silently fail or error)
    - Not checking if `tool_calls` is defined before accessing arguments
    - Using `width`/`height` with FLUX schnell/Kontext models (use `aspect_ratio` instead)
    
    **Common Mistakes:**
    
    - Using OpenAI model names (e.g., `gpt-4o`) with the Together AI SDK -- Together uses Hugging Face-style IDs like `meta-llama/Llama-3.3-70B-Instruct-Turbo`
    - Confusing `client.images.generate()` (Together) with `client.images.create()` (OpenAI) -- different method name
    - Forgetting to use `z.toJSONSchema()` (Zod v4) or `zodToJsonSchema()` (Zod v3) to convert schemas before passing to `response_format`
    - Using the `developer` role (OpenAI-specific) instead of `system` role with Together AI models
    - Passing `max_completion_tokens` instead of `max_tokens` -- Together uses `max_tokens`
    
    **Gotchas & Edge Cases:**
    
    - The SDK auto-retries on 429 (rate limit), 408, 409, and 5xx errors -- 2 retries by default. Disable with `maxRetries: 0`.
    - Model IDs are case-sensitive and follow the `org/model-name` format from Hugging Face.
    - Not all models support function calling. See [examples/tools.md](examples/tools.md) for the current supported list, or check the [official docs](https://docs.together.ai/docs/function-calling).
    - FLUX.1 schnell and Kontext models use `aspect_ratio` parameter; FLUX.1 Pro and FLUX.1.1 Pro use `width`/`height`.
    - Image generation returns URLs by default. Use `response_format: "base64"` for inline data.
    - The `response_format: { type: "json_schema" }` requires telling the model to "only answer in JSON" in the system prompt -- the schema alone is not sufficient.
    - Structured output uses `z.toJSONSchema()` (Zod v4) -- if using Zod v3, use `zodToJsonSchema()` from the `zod-to-json-schema` package.
    - Together AI's `client.images.generate()` is the method name, not `client.images.create()` like OpenAI.
    - Fine-tuning supports LoRA and full fine-tuning. File format is JSONL with `messages` array per line.
    - The OpenAI compatibility endpoint (`api.together.xyz/v1`) supports chat, embeddings, images, vision, function calling, and structured output -- but not fine-tuning or model management.
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST use the `together-ai` package (`import Together from "together-ai"`) -- NOT the OpenAI SDK -- unless explicitly building an OpenAI-compatible integration)**
    
    **(You MUST include the JSON schema in BOTH the `response_format` parameter AND the system prompt when using structured output -- the model needs both)**
    
    **(You MUST handle errors using `Together.APIError` and its subclasses -- never use bare catch blocks without error type checking)**
    
    **(You MUST never hardcode API keys -- always use environment variables via `process.env.TOGETHER_API_KEY`)**
    
    **Failure to follow these rules will produce insecure, unreliable, or incorrectly structured AI integrations.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related