ai-infrastructure-together-ai
Together AI SDK patterns for TypeScript — client setup, chat completions, streaming, structured output, function calling, embeddings, image generation, fine-tuning, and OpenAI-compatible endpoints
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-infrastructure-together-ai/skills/ai-infrastructure-together-ai
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Together AI SDK Patterns
Quick Guide: Use the
together-ainpm package to access 200+ open-source models (Llama, Qwen, Mistral, DeepSeek) via Together AI's fast inference API. The SDK mirrors the OpenAI API shape --client.chat.completions.create()for chat,client.images.generate()for images,client.embeddings.create()for embeddings. Useresponse_format: { type: "json_schema" }with Zod-generated schemas for structured output. Function calling uses the sametoolsparameter shape as OpenAI. You can also use the OpenAI SDK directly by pointingbaseURLtohttps://api.together.xyz/v1.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST use the together-ai package (import Together from "together-ai") -- NOT the OpenAI SDK -- unless explicitly building an OpenAI-compatible integration)
(You MUST include the JSON schema in BOTH the response_format parameter AND the system prompt when using structured output -- the model needs both)
(You MUST handle errors using Together.APIError and its subclasses -- never use bare catch blocks without error type checking)
(You MUST never hardcode API keys -- always use environment variables via process.env.TOGETHER_API_KEY)
</critical_requirements>
Auto-detection: Together AI, together-ai, together.ai, TOGETHER_API_KEY, client.chat.completions (together), client.images.generate, client.embeddings.create (together), Llama-3, Qwen3, Mistral, DeepSeek, FLUX, together.images, together.chat, together.embeddings, together.fineTuning, api.together.xyz
When to use:
- Running open-source LLMs (Llama, Qwen, Mistral, DeepSeek) via serverless inference
- Generating images with FLUX or Stable Diffusion models
- Creating embeddings for RAG pipelines with open-source embedding models
- Using function calling / tool use with open-source models
- Extracting structured JSON output from LLM responses
- Fine-tuning open-source models on custom data
- Migrating from OpenAI to open-source models with minimal code changes
Key patterns covered:
- Client initialization and configuration (retries, timeouts, logging)
- Chat completions with open-source models (Llama, Qwen, Mistral, DeepSeek)
- Streaming with
stream: trueandfor await...of - Structured output with
response_format: { type: "json_schema" }and Zod - Function calling / tool use with
toolsparameter - Image generation with FLUX and Stable Diffusion models
- Embeddings API with open-source embedding models
- Fine-tuning API (file upload, job creation, monitoring)
- OpenAI SDK compatibility (base URL swap)
- Error handling, retries, timeouts
When NOT to use:
- You need OpenAI-specific features (Responses API, Batch API, Realtime API) -- use the OpenAI SDK directly
- You want framework-specific chat UI hooks -- use a framework-integrated AI SDK
- You only use OpenAI models and never plan to use open-source models
Examples Index
- Core: Setup & Configuration -- Client init, production config, error handling, OpenAI compatibility
- Chat Completions -- Basic chat, multi-turn, model selection, vision
- Streaming -- Async iteration, stream cancellation
- Tool/Function Calling -- Tool definitions, multi-step tool loops
- Structured Output -- JSON mode, Zod schemas, regex mode
- Images & Embeddings -- FLUX image generation, embedding models, semantic search
- Quick API Reference -- Model IDs, method signatures, error types
<decision_framework>
Decision Framework
Which Model to Choose
What is your task?
+-- General chat / instruction following -> Llama 3.3 70B Turbo (fast, cheap)
+-- Most capable reasoning -> DeepSeek V3.1, Qwen3.5 397B
+-- Complex math / chain-of-thought -> DeepSeek R1
+-- Function calling / tool use -> Llama 3.3 70B, Qwen3.5 9B
+-- Structured JSON output -> Qwen3.5 9B (best JSON mode support)
+-- Vision / image understanding -> Qwen3-VL-8B-Instruct
+-- Code generation -> DeepSeek V3, Qwen Coder
+-- Embeddings -> BAAI/bge-large-en-v1.5 (default)
+-- Image generation (fast) -> FLUX.1 schnell
+-- Image generation (quality) -> FLUX.2 pro, FLUX.1.1 pro
Together AI SDK vs OpenAI SDK
Do you ONLY use Together AI models?
+-- YES -> Use together-ai package (purpose-built, full API coverage)
+-- NO -> Do you also use OpenAI models?
+-- YES -> Two options:
| +-- Separate SDKs: together-ai for Together, openai for OpenAI
| +-- OpenAI SDK only: Point baseURL to api.together.xyz/v1
+-- NO -> Use a provider-agnostic SDK
Streaming vs Non-Streaming
Is the response user-facing?
+-- YES -> Use streaming (stream: true)
+-- NO -> Use non-streaming
+-- Background processing -> client.chat.completions.create()
+-- Structured output -> Non-streaming with response_format
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Hardcoding
TOGETHER_API_KEYinstead of using environment variables (security breach risk) - Using bare
catchblocks without checkingTogether.APIError(hides API errors) - Not consuming streams returned by
stream: true(tokens are silently lost) - Using
JSON.parse()on completion content withoutresponse_format(fragile, model may return non-JSON) - Omitting the schema from the system prompt when using
response_format: { type: "json_schema" }(degrades output quality)
Medium Priority Issues:
- Not setting
maxRetries/timeoutfor production deployments (default timeout is 1 minute) - Missing
systemrole message (no system instruction means unpredictable behavior) - Using a model that does not support function calling with
toolsparameter (will silently fail or error) - Not checking if
tool_callsis defined before accessing arguments - Using
width/heightwith FLUX schnell/Kontext models (useaspect_ratioinstead)
Common Mistakes:
- Using OpenAI model names (e.g.,
gpt-4o) with the Together AI SDK -- Together uses Hugging Face-style IDs likemeta-llama/Llama-3.3-70B-Instruct-Turbo - Confusing
client.images.generate()(Together) withclient.images.create()(OpenAI) -- different method name - Forgetting to use
z.toJSONSchema()(Zod v4) orzodToJsonSchema()(Zod v3) to convert schemas before passing toresponse_format - Using the
developerrole (OpenAI-specific) instead ofsystemrole with Together AI models - Passing
max_completion_tokensinstead ofmax_tokens-- Together usesmax_tokens
Gotchas & Edge Cases:
- The SDK auto-retries on 429 (rate limit), 408, 409, and 5xx errors -- 2 retries by default. Disable with
maxRetries: 0. - Model IDs are case-sensitive and follow the
org/model-nameformat from Hugging Face. - Not all models support function calling. See examples/tools.md for the current supported list, or check the official docs.
- FLUX.1 schnell and Kontext models use
aspect_ratioparameter; FLUX.1 Pro and FLUX.1.1 Pro usewidth/height. - Image generation returns URLs by default. Use
response_format: "base64"for inline data. - The
response_format: { type: "json_schema" }requires telling the model to "only answer in JSON" in the system prompt -- the schema alone is not sufficient. - Structured output uses
z.toJSONSchema()(Zod v4) -- if using Zod v3, usezodToJsonSchema()from thezod-to-json-schemapackage. - Together AI's
client.images.generate()is the method name, notclient.images.create()like OpenAI. - Fine-tuning supports LoRA and full fine-tuning. File format is JSONL with
messagesarray per line. - The OpenAI compatibility endpoint (
api.together.xyz/v1) supports chat, embeddings, images, vision, function calling, and structured output -- but not fine-tuning or model management.
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST use the together-ai package (import Together from "together-ai") -- NOT the OpenAI SDK -- unless explicitly building an OpenAI-compatible integration)
(You MUST include the JSON schema in BOTH the response_format parameter AND the system prompt when using structured output -- the model needs both)
(You MUST handle errors using Together.APIError and its subclasses -- never use bare catch blocks without error type checking)
(You MUST never hardcode API keys -- always use environment variables via process.env.TOGETHER_API_KEY)
Failure to follow these rules will produce insecure, unreliable, or incorrectly structured AI integrations.
</critical_reminders>
Files (skills)
-
examples
-
chat.md 3.9 KB
# Together AI SDK -- Chat Completions Examples > Chat Completions API patterns: basic completion, multi-turn conversations, model selection, vision, and token control. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [images.md](images.md) -- Image generation, embeddings --- ## Basic Chat Completion ```typescript // basic-chat.ts import Together from "together-ai"; const client = new Together(); async function chat(userMessage: string): Promise<string> { const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [ { role: "system", content: "You are a helpful assistant. Be concise.", }, { role: "user", content: userMessage }, ], }); const content = completion.choices[0].message.content; if (!content) { throw new Error("No content in response"); } return content; } const answer = await chat("What is TypeScript in one sentence?"); console.log(answer); ``` --- ## Multi-Turn Conversations ```typescript import Together from "together-ai"; import type { Together as TogetherTypes } from "together-ai"; const client = new Together(); const messages: TogetherTypes.Chat.CompletionCreateParams["messages"] = [ { role: "system", content: "You are a TypeScript expert." }, { role: "user", content: "What is a union type?" }, ]; const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages, }); // Append assistant response for next turn const assistantMessage = completion.choices[0].message; messages.push({ role: "assistant", content: assistantMessage.content ?? "" }); messages.push({ role: "user", content: "Give me a real-world example." }); const followUp = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages, }); ``` --- ## Controlling Output Length and Temperature ```typescript const MAX_TOKENS = 500; const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Summarize this article." }], max_tokens: MAX_TOKENS, temperature: 0, // Deterministic output }); const finishReason = completion.choices[0].finish_reason; if (finishReason === "length") { console.warn("Output was truncated -- increase max_tokens"); } ``` --- ## Vision -- Image from URL ```typescript import Together from "together-ai"; const client = new Together(); async function analyzeImage( imageUrl: string, question: string, ): Promise<string> { const response = await client.chat.completions.create({ model: "Qwen/Qwen3-VL-8B-Instruct", messages: [ { role: "user", content: [ { type: "text", text: question }, { type: "image_url", image_url: { url: imageUrl } }, ], }, ], }); return response.choices[0].message.content ?? ""; } const description = await analyzeImage( "https://example.com/photo.jpg", "Describe what you see in this image.", ); console.log(description); ``` --- ## Vision -- Multiple Images ```typescript const response = await client.chat.completions.create({ model: "Qwen/Qwen3-VL-8B-Instruct", messages: [ { role: "user", content: [ { type: "text", text: "Compare these two images." }, { type: "image_url", image_url: { url: "https://example.com/image1.jpg" }, }, { type: "image_url", image_url: { url: "https://example.com/image2.jpg" }, }, ], }, ], }); ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
core.md 5.7 KB
# Together AI SDK -- Setup & Configuration Examples > Client initialization, environment config, production settings, error handling, and OpenAI compatibility. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [images.md](images.md) -- Image generation, embeddings --- ## Basic Client Setup ```typescript // lib/together.ts import Together from "together-ai"; // Reads TOGETHER_API_KEY from env automatically const client = new Together(); export { client }; ``` --- ## Production Configuration ```typescript // lib/together.ts import Together from "together-ai"; const TIMEOUT_MS = 30_000; const MAX_RETRIES = 3; const client = new Together({ apiKey: process.env.TOGETHER_API_KEY, timeout: TIMEOUT_MS, maxRetries: MAX_RETRIES, }); export { client }; ``` --- ## Production Error Handling ```typescript // error-handling.ts import Together from "together-ai"; const TIMEOUT_MS = 30_000; const MAX_RETRIES = 3; const client = new Together({ timeout: TIMEOUT_MS, maxRetries: MAX_RETRIES, }); async function safeCompletion(prompt: string): Promise<string | null> { try { const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: prompt }, ], }); const content = completion.choices[0].message.content; if (!content) { throw new Error("No content in response"); } return content; } catch (error) { if (error instanceof Together.APIError) { console.error(`Together API Error [${error.status}]: ${error.message}`); if (error instanceof Together.RateLimitError) { console.error("Rate limited. SDK will auto-retry."); // If we get here, all retries were exhausted return null; } if (error instanceof Together.AuthenticationError) { throw new Error( "Invalid API key. Check TOGETHER_API_KEY environment variable.", ); } if (error instanceof Together.BadRequestError) { console.error("Invalid request parameters:", error.message); return null; } if (error instanceof Together.InternalServerError) { console.error("Together server error after all retries"); return null; } } // Network/connection errors if (error instanceof Together.APIConnectionError) { console.error("Network error:", error.message); return null; } // Unknown errors should be re-thrown throw error; } } const result = await safeCompletion("Hello!"); if (result) { console.log(result); } else { console.error("Failed to get completion"); } ``` --- ## Error Type Hierarchy ```typescript // Error class hierarchy: // Together.APIError (base) // +-- Together.BadRequestError (400) // +-- Together.AuthenticationError (401) // +-- Together.PermissionDeniedError (403) // +-- Together.NotFoundError (404) // +-- Together.UnprocessableEntityError (422) // +-- Together.RateLimitError (429) // +-- Together.InternalServerError (>=500) // +-- Together.APIConnectionError (network) ``` --- ## OpenAI SDK Compatibility Use the OpenAI SDK with Together AI by swapping the base URL. ```typescript // lib/together-openai.ts import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.TOGETHER_API_KEY, baseURL: "https://api.together.xyz/v1", }); // Use exactly like OpenAI SDK, but with Together model IDs const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: "Hello!" }, ], }); console.log(completion.choices[0].message.content); export { client }; ``` --- ## Per-Request Overrides ```typescript // Override retries and timeout for a single request await client.chat.completions.create( { model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], }, { maxRetries: 5, timeout: 60_000, signal: abortController.signal, headers: { "X-Custom-Header": "value" }, }, ); ``` --- ## Request Cancellation with AbortController ```typescript const controller = new AbortController(); const ABORT_TIMEOUT_MS = 5_000; // Cancel after timeout setTimeout(() => controller.abort(), ABORT_TIMEOUT_MS); try { const completion = await client.chat.completions.create( { model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], }, { signal: controller.signal }, ); } catch (error) { if (error instanceof Error && error.name === "AbortError") { console.log("Request was cancelled"); } } ``` --- ## Raw Response Access ```typescript // Get underlying Response object const response = await client.chat.completions .create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], }) .asResponse(); console.log(response.status); console.log(response.headers.get("x-request-id")); // Or get both data and response const { data, response: raw } = await client.chat.completions .create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], }) .withResponse(); ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
images.md 7.4 KB
# Together AI SDK -- Images & Embeddings Examples > Image generation with FLUX/Stable Diffusion and embeddings for semantic search. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod --- ## Basic Image Generation (FLUX schnell) ```typescript // image-generation.ts import Together from "together-ai"; const client = new Together(); const response = await client.images.generate({ model: "black-forest-labs/FLUX.1-schnell", prompt: "A serene mountain landscape at sunset with a lake reflection", steps: 4, }); console.log(response.data[0].url); ``` --- ## High Quality Image (FLUX Pro) ```typescript import Together from "together-ai"; const client = new Together(); const IMAGE_WIDTH = 1024; const IMAGE_HEIGHT = 768; const IMAGE_STEPS = 25; const response = await client.images.generate({ model: "black-forest-labs/FLUX.1.1-pro", prompt: "A photorealistic portrait of a robot in a garden", width: IMAGE_WIDTH, height: IMAGE_HEIGHT, steps: IMAGE_STEPS, }); console.log(response.data[0].url); ``` --- ## Multiple Image Variations ```typescript import Together from "together-ai"; const client = new Together(); const NUM_VARIATIONS = 4; const response = await client.images.generate({ model: "black-forest-labs/FLUX.1-schnell", prompt: "A cute robot assistant helping in a modern office", n: NUM_VARIATIONS, steps: 4, }); response.data.forEach((image, index) => { console.log(`Variation ${index + 1}: ${image.url}`); }); ``` --- ## Base64 Response Format ```typescript import Together from "together-ai"; const client = new Together(); const response = await client.images.generate({ model: "black-forest-labs/FLUX.1-schnell", prompt: "A cat in outer space", response_format: "base64", }); const base64Data = response.data[0].b64_json; // Use directly in <img src="data:image/png;base64,${base64Data}" /> ``` --- ## Image Editing with Reference Images (FLUX.2) ```typescript import Together from "together-ai"; const client = new Together(); const response = await client.images.generate({ model: "black-forest-labs/FLUX.2-pro", prompt: "Replace the color of the car to blue", width: 1024, height: 768, reference_images: ["https://example.com/original-car.jpg"], }); console.log(response.data[0].url); ``` --- ## Kontext Image Editing ```typescript import Together from "together-ai"; const client = new Together(); const response = await client.images.generate({ model: "black-forest-labs/FLUX.1-kontext-pro", prompt: "Add a party hat to the dog", image_url: "https://example.com/dog.jpg", }); console.log(response.data[0].url); ``` --- ## Image Model Selection Guide ``` Fast generation (4 steps) -> FLUX.1 schnell (aspect_ratio param) High quality -> FLUX.1.1 pro (width/height params) Latest + reference images -> FLUX.2 pro (width/height + reference_images) Image editing (single ref) -> FLUX.1 Kontext pro (aspect_ratio + image_url) Negative prompts -> Stable Diffusion models (negative_prompt param) ``` --- ## Basic Embeddings ```typescript // embeddings.ts import Together from "together-ai"; const client = new Together(); const EMBEDDING_MODEL = "BAAI/bge-large-en-v1.5"; const response = await client.embeddings.create({ model: EMBEDDING_MODEL, input: "TypeScript provides static type checking for JavaScript.", }); console.log("Embedding dimensions:", response.data[0].embedding.length); console.log("First 5 values:", response.data[0].embedding.slice(0, 5)); ``` --- ## Batch Embeddings ```typescript import Together from "together-ai"; const client = new Together(); const EMBEDDING_MODEL = "BAAI/bge-large-en-v1.5"; const documents = [ "TypeScript provides static type checking.", "React is a library for building user interfaces.", "Node.js is a JavaScript runtime built on V8.", "PostgreSQL is a powerful relational database.", ]; // Batch all inputs in one call for efficiency const response = await client.embeddings.create({ model: EMBEDDING_MODEL, input: documents, }); const embeddings = response.data.map((item) => ({ index: item.index, embedding: item.embedding, })); console.log(`Generated ${embeddings.length} embeddings`); ``` --- ## Semantic Search with Cosine Similarity ```typescript import Together from "together-ai"; const client = new Together(); const EMBEDDING_MODEL = "BAAI/bge-large-en-v1.5"; const SIMILARITY_THRESHOLD = 0.7; const TOP_K = 3; function cosineSimilarity(a: number[], b: number[]): number { let dot = 0; let normA = 0; let normB = 0; for (let i = 0; i < a.length; i++) { dot += a[i] * b[i]; normA += a[i] * a[i]; normB += b[i] * b[i]; } return dot / (Math.sqrt(normA) * Math.sqrt(normB)); } // Index documents const documents = [ "TypeScript provides static type checking for JavaScript.", "React is a library for building user interfaces.", "PostgreSQL is a powerful relational database.", "Docker containers package applications with dependencies.", ]; const docEmbeddings = await client.embeddings.create({ model: EMBEDDING_MODEL, input: documents, }); const indexedDocs = documents.map((text, i) => ({ text, embedding: docEmbeddings.data[i].embedding, })); // Search async function search( query: string, ): Promise<Array<{ text: string; score: number }>> { const queryEmbedding = await client.embeddings.create({ model: EMBEDDING_MODEL, input: query, }); const queryVector = queryEmbedding.data[0].embedding; return indexedDocs .map((doc) => ({ text: doc.text, score: cosineSimilarity(queryVector, doc.embedding), })) .filter((r) => r.score > SIMILARITY_THRESHOLD) .sort((a, b) => b.score - a.score) .slice(0, TOP_K); } const results = await search("What is TypeScript?"); results.forEach((r) => { console.log(`[${r.score.toFixed(3)}] ${r.text}`); }); ``` --- ## Fine-Tuning (File Upload + Job Creation) ```typescript import Together from "together-ai"; import { createReadStream } from "node:fs"; const client = new Together(); const EPOCHS = 3; // Upload training data (JSONL format) const file = await client.files.upload({ file: createReadStream("training-data.jsonl"), purpose: "fine-tune", }); console.log(`Uploaded file: ${file.id}`); // Create fine-tuning job const job = await client.fineTuning.create({ training_file: file.id, model: "meta-llama/Meta-Llama-3-8B-Instruct", n_epochs: EPOCHS, }); console.log(`Fine-tuning job: ${job.id}`); // Monitor job status const status = await client.fineTuning.retrieve(job.id); console.log(`Status: ${status.status}`); // List events const events = await client.fineTuning.listEvents(job.id); console.log("Events:", events); ``` ### Training Data Format (JSONL) Each line is a JSON object with a `messages` array: ```jsonl {"messages": [{"role": "system", "content": "You are helpful."}, {"role": "user", "content": "Hi"}, {"role": "assistant", "content": "Hello!"}]} {"messages": [{"role": "user", "content": "What is TypeScript?"}, {"role": "assistant", "content": "TypeScript is a typed superset of JavaScript."}]} ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
streaming.md 3.6 KB
# Together AI SDK -- Streaming Examples > Streaming patterns: `stream: true` with async iterators, stream cancellation, and controller access. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat Completions API - [tools.md](tools.md) -- Tool/function calling - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [images.md](images.md) -- Image generation, embeddings --- ## Basic Streaming with `for await` ```typescript // streaming-chat.ts import Together from "together-ai"; const client = new Together(); const stream = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: "Explain async/await in TypeScript." }, ], stream: true, }); for await (const chunk of stream) { const content = chunk.choices[0]?.delta?.content; if (content) { process.stdout.write(content); } } console.log(); // newline ``` --- ## Collecting Full Response from Stream ```typescript import Together from "together-ai"; const client = new Together(); async function streamToString(prompt: string): Promise<string> { const stream = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: prompt }, ], stream: true, }); const parts: string[] = []; for await (const chunk of stream) { const content = chunk.choices[0]?.delta?.content; if (content) { parts.push(content); process.stdout.write(content); // Show progress } } console.log(); // newline return parts.join(""); } const result = await streamToString("Explain promises in JavaScript."); console.log("Total length:", result.length); ``` --- ## Stream Cancellation ```typescript import Together from "together-ai"; const client = new Together(); const stream = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Tell me a long story." }], stream: true, }); let tokenCount = 0; const MAX_TOKENS_TO_READ = 100; for await (const chunk of stream) { const content = chunk.choices[0]?.delta?.content; if (content) { process.stdout.write(content); tokenCount++; if (tokenCount >= MAX_TOKENS_TO_READ) { // Cancel the stream by using the controller stream.controller.abort(); break; } } } console.log("\nStream cancelled after", tokenCount, "chunks"); ``` --- ## Streaming with Tool Calls ```typescript import Together from "together-ai"; const client = new Together(); const stream = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "What is the weather in Tokyo?" }], tools: [ { type: "function", function: { name: "get_weather", description: "Get weather for a city", parameters: { type: "object", properties: { location: { type: "string" }, }, required: ["location"], }, }, }, ], stream: true, }); for await (const chunk of stream) { const toolCalls = chunk.choices[0]?.delta?.tool_calls ?? []; for (const toolCall of toolCalls) { console.log("Tool call delta:", toolCall); } } ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
structured-output.md 5.4 KB
# Together AI SDK -- Structured Output Examples > Structured output patterns: JSON schema mode with Zod, regex mode, vision with JSON, and complex schemas. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [tools.md](tools.md) -- Tool/function calling - [images.md](images.md) -- Image generation, embeddings --- ## JSON Schema Mode with Zod ```typescript // structured-output.ts import Together from "together-ai"; import { z } from "zod"; const client = new Together(); const VoiceNoteSchema = z.object({ title: z.string().describe("A title for the voice note"), summary: z.string().describe("A one sentence summary"), actionItems: z .array(z.string()) .describe("A list of action items from the note"), }); type VoiceNote = z.infer<typeof VoiceNoteSchema>; const jsonSchema = z.toJSONSchema(VoiceNoteSchema); async function extractVoiceNote(transcript: string): Promise<VoiceNote> { const completion = await client.chat.completions.create({ model: "Qwen/Qwen3.5-9B", messages: [ { role: "system", content: `Extract structured data from the transcript. Only answer in JSON. Follow this schema: ${JSON.stringify(jsonSchema)}`, }, { role: "user", content: transcript }, ], response_format: { type: "json_schema", json_schema: { name: "voice_note", schema: jsonSchema, }, }, }); return JSON.parse(completion.choices[0].message.content ?? "{}"); } const note = await extractVoiceNote( "Need to buy groceries today and schedule a meeting with the team for Friday.", ); console.log(note.title); console.log(note.actionItems); ``` --- ## Complex Nested Schema ```typescript import Together from "together-ai"; import { z } from "zod"; const client = new Together(); const ArticleSummary = z.object({ title: z.string(), summary: z.string(), keyPoints: z.array(z.string()), sentiment: z.enum(["positive", "negative", "neutral"]), topics: z.array( z.object({ name: z.string(), relevance: z.number().describe("Relevance score 0-1"), }), ), }); const jsonSchema = z.toJSONSchema(ArticleSummary); const completion = await client.chat.completions.create({ model: "Qwen/Qwen3.5-9B", messages: [ { role: "system", content: `Extract article summary. Only answer in JSON. Schema: ${JSON.stringify(jsonSchema)}`, }, { role: "user", content: "TypeScript 5.8 brings improved type inference and better error messages, making everyday coding more productive.", }, ], response_format: { type: "json_schema", json_schema: { name: "article_summary", schema: jsonSchema }, }, }); const article = JSON.parse(completion.choices[0].message.content ?? "{}"); console.log(`Title: ${article.title}`); console.log(`Sentiment: ${article.sentiment}`); ``` --- ## Regex Mode (Constrained Output) Constrain output to a regex pattern for classification tasks. ```typescript import Together from "together-ai"; const client = new Together(); const MAX_TOKENS = 10; const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", temperature: 0.2, max_tokens: MAX_TOKENS, messages: [ { role: "system", content: "Classify the sentiment of the text as positive, neutral, or negative.", }, { role: "user", content: "Wow. I loved the movie!" }, ], response_format: { type: "regex", // @ts-ignore -- regex type not in SDK types yet pattern: "(positive|neutral|negative)", }, }); console.log(completion.choices[0].message.content); // Output: "positive" ``` --- ## Vision Model with JSON Output Extract structured data from images. ```typescript import Together from "together-ai"; import { z } from "zod"; const client = new Together(); const ImageDescription = z.object({ description: z.string().describe("What the image shows"), objectCount: z.number().describe("Number of main objects"), dominantColors: z.array(z.string()).describe("Main colors in the image"), }); const jsonSchema = z.toJSONSchema(ImageDescription); const completion = await client.chat.completions.create({ model: "Qwen/Qwen3-VL-8B-Instruct", messages: [ { role: "user", content: [ { type: "text", text: `Describe this image. Only answer in JSON. Schema: ${JSON.stringify(jsonSchema)}`, }, { type: "image_url", image_url: { url: "https://example.com/photo.jpg" }, }, ], }, ], response_format: { type: "json_schema", json_schema: { name: "image_description", schema: jsonSchema }, }, }); const description = JSON.parse(completion.choices[0].message.content ?? "{}"); console.log(description); ``` --- ## Zod v3 vs Zod v4 ```typescript // Zod v4 (recommended): Use built-in z.toJSONSchema() import { z } from "zod"; const schema = z.object({ name: z.string() }); const jsonSchema = z.toJSONSchema(schema); // Built-in // Zod v3: Use zodToJsonSchema from separate package import { z } from "zod"; import { zodToJsonSchema } from "zod-to-json-schema"; const schema = z.object({ name: z.string() }); const jsonSchema = zodToJsonSchema(schema); // External package ``` --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._ -
tools.md 7.6 KB
# Together AI SDK -- Tool/Function Calling Examples > Function calling patterns: tool definitions, tool_choice control, multi-step tool loops, parallel calls. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [core.md](core.md) -- Client setup, error handling - [chat.md](chat.md) -- Chat Completions API - [streaming.md](streaming.md) -- Streaming responses - [structured-output.md](structured-output.md) -- Structured outputs with Zod - [images.md](images.md) -- Image generation, embeddings --- ## Basic Function Calling ```typescript import Together from "together-ai"; const client = new Together(); const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "What is the weather in Tokyo?" }], tools: [ { type: "function", function: { name: "get_weather", description: "Get the current weather for a location", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, unit: { type: "string", enum: ["celsius", "fahrenheit"], description: "Temperature unit", }, }, required: ["location"], additionalProperties: false, }, strict: true, }, }, ], }); const toolCall = completion.choices[0].message.tool_calls?.[0]; if (toolCall) { const args = JSON.parse(toolCall.function.arguments); console.log(`Call ${toolCall.function.name} with:`, args); } ``` --- ## Multiple Tools ```typescript import Together from "together-ai"; const client = new Together(); const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [ { role: "system", content: "You help users by calling available tools.", }, { role: "user", content: "What is the weather in London and search for restaurants?", }, ], tools: [ { type: "function", function: { name: "get_weather", description: "Get current weather for a city", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, }, required: ["location"], additionalProperties: false, }, }, }, { type: "function", function: { name: "search_restaurants", description: "Search for restaurants in a city", parameters: { type: "object", properties: { city: { type: "string", description: "City to search" }, cuisine: { type: "string", description: "Cuisine type" }, }, required: ["city"], additionalProperties: false, }, }, }, ], }); // Model may return one or more tool calls const toolCalls = completion.choices[0].message.tool_calls ?? []; for (const toolCall of toolCalls) { const args = JSON.parse(toolCall.function.arguments); console.log(`Call ${toolCall.function.name}:`, args); } ``` --- ## Controlling Tool Invocation with `tool_choice` ```typescript // Force a specific function const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Tell me about Paris" }], tools: [ /* ... */ ], tool_choice: { type: "function", function: { name: "get_weather" }, }, }); // Force at least one tool call (any tool) const required = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], tools: [ /* ... */ ], tool_choice: "required", }); // Disable tool calling for this request const noTools = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], tools: [ /* ... */ ], tool_choice: "none", }); ``` --- ## Multi-Step Tool Loop Execute tool calls and feed results back to the model. ```typescript import Together from "together-ai"; const client = new Together(); // Tool implementations async function getWeather(location: string): Promise<string> { return JSON.stringify({ location, temperature: 22, condition: "sunny" }); } async function searchDatabase(query: string): Promise<string> { return JSON.stringify({ results: [{ id: 1, title: `Result: ${query}` }] }); } const toolImplementations: Record< string, (args: Record<string, string>) => Promise<string> > = { get_weather: (args) => getWeather(args.location), search_database: (args) => searchDatabase(args.query), }; const tools = [ { type: "function" as const, function: { name: "get_weather", description: "Get weather for a city", parameters: { type: "object" as const, properties: { location: { type: "string", description: "City name" }, }, required: ["location"], additionalProperties: false, }, }, }, { type: "function" as const, function: { name: "search_database", description: "Search the database", parameters: { type: "object" as const, properties: { query: { type: "string", description: "Search query" }, }, required: ["query"], additionalProperties: false, }, }, }, ]; // Initial request const messages: Array<{ role: "system" | "user" | "assistant" | "tool"; content: string | null; tool_calls?: Array<{ id: string; type: "function"; function: { name: string; arguments: string }; }>; tool_call_id?: string; name?: string; }> = [ { role: "system", content: "Help users with weather and database queries." }, { role: "user", content: "What is the weather in London?" }, ]; const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages, tools, }); const assistantMessage = completion.choices[0].message; // If model wants to call tools, execute them and continue if (assistantMessage.tool_calls && assistantMessage.tool_calls.length > 0) { // Add assistant message with tool calls messages.push({ role: "assistant", content: assistantMessage.content, tool_calls: assistantMessage.tool_calls, }); // Execute each tool and add results for (const toolCall of assistantMessage.tool_calls) { const args = JSON.parse(toolCall.function.arguments); const impl = toolImplementations[toolCall.function.name]; const result = impl ? await impl(args) : "Unknown tool"; messages.push({ role: "tool", content: result, tool_call_id: toolCall.id, name: toolCall.function.name, }); } // Get final response without tools const finalCompletion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages, }); console.log("Final answer:", finalCompletion.choices[0].message.content); } ``` --- ## Supported Models for Function Calling Not all Together AI models support function calling. Currently supported: - `meta-llama/Llama-3.3-70B-Instruct-Turbo` - `meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8` - `Qwen/Qwen3.5-397B-A17B` - `Qwen/Qwen3.5-9B` - `Qwen/Qwen3-Next-80B-A3B-Instruct` - `Qwen/Qwen2.5-7B-Instruct-Turbo` - `deepseek-ai/DeepSeek-V3` - `deepseek-ai/DeepSeek-R1` - `mistralai/Mistral-Small-24B-Instruct-2501` Check [Together AI docs](https://docs.together.ai/docs/function-calling) for the latest model support. --- _For core concepts, see [SKILL.md](../SKILL.md). For API reference tables, see [reference.md](../reference.md)._
-
-
reference.md 8.4 KB
# Together AI SDK Quick Reference > Client configuration, model IDs, API methods, error types, and image parameters. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Package Installation ```bash # Core package (always required) npm install together-ai # For structured outputs (optional but recommended) npm install zod ``` --- ## Client Configuration ```typescript import Together from "together-ai"; const client = new Together({ apiKey: process.env.TOGETHER_API_KEY, // Auto-reads from env if not set timeout: 30_000, // Request timeout in ms (default: 60_000 = 1 min) maxRetries: 3, // Retry count on 429/5xx (default: 2) baseURL: "https://api.together.xyz/v1", // Override for proxies logLevel: "off", // 'debug' | 'info' | 'warn' | 'error' | 'off' }); ``` ### Environment Variables | Variable | Purpose | | ------------------ | ----------------------- | | `TOGETHER_API_KEY` | API key (auto-detected) | --- ## Model IDs ### Chat / Language Models | Model ID | Use Case | Notes | | --------------------------------------------------- | --------------------------- | ----------------------------- | | `meta-llama/Llama-3.3-70B-Instruct-Turbo` | General purpose, fast | Best balance of speed/quality | | `meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8` | Latest Llama 4 | MoE architecture | | `Qwen/Qwen3.5-9B` | JSON mode, function calling | Excellent structured output | | `Qwen/Qwen3.5-397B-A17B` | Most capable Qwen | MoE, hybrid reasoning | | `deepseek-ai/DeepSeek-V3.1` | Most capable open-source | Strong reasoning | | `deepseek-ai/DeepSeek-R1` | Complex reasoning | Chain-of-thought reasoning | | `mistralai/Mistral-Small-24B-Instruct-2501` | Fast, function calling | Good tool use support | | `google/gemma-3n-E4B-it` | Ultra-lightweight | Cheapest option | ### Vision / Multimodal Models | Model ID | Use Case | | ------------------------------------------------ | ------------------- | | `Qwen/Qwen3-VL-8B-Instruct` | Image understanding | | `meta-llama/Llama-3.2-11B-Vision-Instruct-Turbo` | Vision + chat | ### Embedding Models | Model ID | Use Case | | ------------------------------------------- | ---------------------- | | `BAAI/bge-large-en-v1.5` | General English | | `WhereIsAI/UAE-Large-V1` | High quality | | `intfloat/multilingual-e5-large-instruct` | Multilingual | | `togethercomputer/m2-bert-80M-8k-retrieval` | Long context retrieval | ### Image Generation Models | Model ID | Use Case | Size Control | | -------------------------------------- | ---------------------- | ------------------ | | `black-forest-labs/FLUX.1-schnell` | Fast generation | `aspect_ratio` | | `black-forest-labs/FLUX.1.1-pro` | High quality | `width` / `height` | | `black-forest-labs/FLUX.2-pro` | Latest, reference imgs | `width` / `height` | | `black-forest-labs/FLUX.1-kontext-pro` | Image editing | `aspect_ratio` | --- ## API Methods Reference ### Chat Completions ```typescript const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", // Required messages: [], // Required: ChatCompletionMessageParam[] temperature: 0.7, // 0-2 (default: 1) max_tokens: 1000, // Max output tokens top_p: 1, // Nucleus sampling tools: [], // Function calling tools tool_choice: "auto", // 'auto' | 'required' | 'none' | { type: 'function', function: { name } } response_format: undefined, // { type: 'json_schema', json_schema: { name, schema } } stream: false, // Enable streaming stop: undefined, // string | string[] -- stop sequences }); ``` ### Images ```typescript const response = await client.images.generate({ model: "black-forest-labs/FLUX.1-schnell", // Required prompt: "", // Required (except Kling) width: 1024, // For FLUX Pro/1.1 Pro/Dev height: 1024, // For FLUX Pro/1.1 Pro/Dev aspect_ratio: undefined, // For FLUX schnell/Kontext: "1:1", "16:9", etc. steps: 4, // 1-50 n: 1, // 1-4 variations seed: undefined, // Reproducibility negative_prompt: "", // Unwanted elements response_format: undefined, // "base64" for inline data reference_images: [], // FLUX.2, Google models image_url: undefined, // Kontext single reference disable_safety_checker: false, // NSFW filter toggle }); ``` ### Embeddings ```typescript const response = await client.embeddings.create({ model: "BAAI/bge-large-en-v1.5", // Required input: "", // string or string[] }); ``` ### Fine-Tuning ```typescript // Upload training data const file = await client.files.upload({ file: readStream, // ReadStream or File purpose: "fine-tune", }); // Create fine-tuning job const job = await client.fineTuning.create({ training_file: file.id, // Required model: "meta-llama/Meta-Llama-3-8B-Instruct", // Required n_epochs: 3, }); // Monitor job const status = await client.fineTuning.retrieve(job.id); const events = await client.fineTuning.listEvents(job.id); // List, cancel, delete const jobs = await client.fineTuning.list(); await client.fineTuning.cancel(job.id); await client.fineTuning.delete(job.id); ``` ### Files ```typescript // Upload const file = await client.files.upload({ file: readStream, // ReadStream or File purpose: "fine-tune", // 'fine-tune' }); // List / Retrieve / Delete const files = await client.files.list(); const fileInfo = await client.files.retrieve("file-abc123"); await client.files.delete("file-abc123"); ``` --- ## Error Types | Error Class | HTTP Status | Auto-Retried? | | --------------------------- | ----------- | ------------- | | `BadRequestError` | 400 | No | | `AuthenticationError` | 401 | No | | `PermissionDeniedError` | 403 | No | | `NotFoundError` | 404 | No | | `UnprocessableEntityError` | 422 | No | | `RateLimitError` | 429 | Yes | | `InternalServerError` | >= 500 | Yes | | `APIConnectionError` | N/A | Yes | | `APIConnectionTimeoutError` | N/A | Yes | All errors extend `Together.APIError` with properties: - `.status` -- HTTP status code - `.message` -- Error message --- ## OpenAI Compatibility ```typescript import OpenAI from "openai"; const client = new OpenAI({ apiKey: process.env.TOGETHER_API_KEY, baseURL: "https://api.together.xyz/v1", }); // Use exactly like OpenAI SDK, with Together AI model IDs const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], }); ``` **Compatible endpoints:** chat completions, embeddings, images, vision, function calling, structured output **NOT compatible:** fine-tuning management, model listing, Together-specific endpoints --- ## Message Roles | Role | Description | Notes | | ----------- | ------------------ | ----------------------------- | | `system` | System instruction | Use `system`, NOT `developer` | | `user` | User input | | | `assistant` | Model response | | | `tool` | Tool result | Used in multi-step tool flows | --- ## Response Format Options | Type | Shape | Use Case | | ------------- | ------------------------------------------------------------- | ----------------------- | | `json_schema` | `{ type: "json_schema", json_schema: { name, schema } }` | Structured JSON output | | `json_object` | `{ type: "json_object" }` | Freeform JSON | | `regex` | `{ type: "regex", pattern: "(positive\|negative\|neutral)" }` | Constrained text output | -
SKILL.md 19.4 KB
--- name: ai-infrastructure-together-ai description: Together AI SDK patterns for TypeScript — client setup, chat completions, streaming, structured output, function calling, embeddings, image generation, fine-tuning, and OpenAI-compatible endpoints --- # Together AI SDK Patterns > **Quick Guide:** Use the `together-ai` npm package to access 200+ open-source models (Llama, Qwen, Mistral, DeepSeek) via Together AI's fast inference API. The SDK mirrors the OpenAI API shape -- `client.chat.completions.create()` for chat, `client.images.generate()` for images, `client.embeddings.create()` for embeddings. Use `response_format: { type: "json_schema" }` with Zod-generated schemas for structured output. Function calling uses the same `tools` parameter shape as OpenAI. You can also use the OpenAI SDK directly by pointing `baseURL` to `https://api.together.xyz/v1`. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST use the `together-ai` package (`import Together from "together-ai"`) -- NOT the OpenAI SDK -- unless explicitly building an OpenAI-compatible integration)** **(You MUST include the JSON schema in BOTH the `response_format` parameter AND the system prompt when using structured output -- the model needs both)** **(You MUST handle errors using `Together.APIError` and its subclasses -- never use bare catch blocks without error type checking)** **(You MUST never hardcode API keys -- always use environment variables via `process.env.TOGETHER_API_KEY`)** </critical_requirements> --- **Auto-detection:** Together AI, together-ai, together.ai, TOGETHER_API_KEY, client.chat.completions (together), client.images.generate, client.embeddings.create (together), Llama-3, Qwen3, Mistral, DeepSeek, FLUX, together.images, together.chat, together.embeddings, together.fineTuning, api.together.xyz **When to use:** - Running open-source LLMs (Llama, Qwen, Mistral, DeepSeek) via serverless inference - Generating images with FLUX or Stable Diffusion models - Creating embeddings for RAG pipelines with open-source embedding models - Using function calling / tool use with open-source models - Extracting structured JSON output from LLM responses - Fine-tuning open-source models on custom data - Migrating from OpenAI to open-source models with minimal code changes **Key patterns covered:** - Client initialization and configuration (retries, timeouts, logging) - Chat completions with open-source models (Llama, Qwen, Mistral, DeepSeek) - Streaming with `stream: true` and `for await...of` - Structured output with `response_format: { type: "json_schema" }` and Zod - Function calling / tool use with `tools` parameter - Image generation with FLUX and Stable Diffusion models - Embeddings API with open-source embedding models - Fine-tuning API (file upload, job creation, monitoring) - OpenAI SDK compatibility (base URL swap) - Error handling, retries, timeouts **When NOT to use:** - You need OpenAI-specific features (Responses API, Batch API, Realtime API) -- use the OpenAI SDK directly - You want framework-specific chat UI hooks -- use a framework-integrated AI SDK - You only use OpenAI models and never plan to use open-source models --- ## Examples Index - [Core: Setup & Configuration](examples/core.md) -- Client init, production config, error handling, OpenAI compatibility - [Chat Completions](examples/chat.md) -- Basic chat, multi-turn, model selection, vision - [Streaming](examples/streaming.md) -- Async iteration, stream cancellation - [Tool/Function Calling](examples/tools.md) -- Tool definitions, multi-step tool loops - [Structured Output](examples/structured-output.md) -- JSON mode, Zod schemas, regex mode - [Images & Embeddings](examples/images.md) -- FLUX image generation, embedding models, semantic search - [Quick API Reference](reference.md) -- Model IDs, method signatures, error types --- <philosophy> ## Philosophy Together AI provides **fast serverless inference for open-source models**. The TypeScript SDK (`together-ai`) is auto-generated with Stainless and mirrors the OpenAI API shape, making migration straightforward. **Core principles:** 1. **OpenAI-compatible API shape** -- Same `client.chat.completions.create()` pattern, same `messages` array, same `tools` parameter. Switching from OpenAI is often just changing the import and model name. 2. **Open-source model access** -- Run Llama, Qwen, Mistral, DeepSeek, and 200+ other models without managing infrastructure. Models are identified by their Hugging Face-style IDs (e.g., `meta-llama/Llama-3.3-70B-Instruct-Turbo`). 3. **Multi-modal support** -- Chat completions, image generation (FLUX, Stable Diffusion), embeddings, audio, and video -- all through one SDK. 4. **Structured output via JSON Schema** -- Pass a JSON schema in `response_format` and include it in the system prompt. Use Zod's `z.toJSONSchema()` to generate schemas from TypeScript types. 5. **Fine-tuning open-source models** -- Upload JSONL data, create LoRA or full fine-tuning jobs, and deploy custom models -- all via the API. **When to use Together AI:** - You want to use open-source models with fast serverless inference - You need cost-effective inference (often cheaper than proprietary APIs) - You want to fine-tune open-source models on your data - You need image generation with FLUX models - You want OpenAI API compatibility for easy migration **When NOT to use:** - You need OpenAI-specific features (Responses API, Batch API, Realtime) -- use the OpenAI SDK - You need Anthropic or Google-specific features -- use their respective SDKs - You want a provider-agnostic SDK -- use a unified provider framework </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Setup Initialize the Together client. It reads `TOGETHER_API_KEY` from the environment. ```typescript // lib/together.ts -- basic setup import Together from "together-ai"; const client = new Together(); export { client }; ``` ```typescript // lib/together.ts -- production configuration const TIMEOUT_MS = 30_000; const MAX_RETRIES = 3; const client = new Together({ apiKey: process.env.TOGETHER_API_KEY, timeout: TIMEOUT_MS, maxRetries: MAX_RETRIES, }); export { client }; ``` **Why good:** Minimal setup, env var auto-detected, named constants for production settings ```typescript // BAD: Hardcoded API key const client = new Together({ apiKey: "sk-abc123...", }); ``` **Why bad:** Hardcoded keys get leaked in version control, security breach risk **See:** [examples/core.md](examples/core.md) for error handling, OpenAI compatibility, per-request overrides --- ### Pattern 2: Chat Completions Stateless text generation with open-source models. ```typescript const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [ { role: "system", content: "You are a helpful coding assistant." }, { role: "user", content: "Explain TypeScript generics." }, ], }); console.log(completion.choices[0].message.content); ``` **Why good:** Clear message roles, system message for behavior control, direct content access ```typescript // BAD: No system message, no model specified const res = await client.chat.completions.create({ messages: [{ role: "user", content: "do something" }], }); ``` **Why bad:** Missing `model` field will error, no system instruction means unpredictable behavior **See:** [examples/chat.md](examples/chat.md) for multi-turn, vision models, model selection guide --- ### Pattern 3: Streaming Use streaming for user-facing responses. ```typescript const stream = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Explain async/await." }], stream: true, }); for await (const chunk of stream) { const content = chunk.choices[0]?.delta?.content; if (content) process.stdout.write(content); } ``` **Why good:** Progressive output for better UX, standard async iterator pattern ```typescript // BAD: Not consuming the stream const stream = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], stream: true, }); // Stream never consumed -- tokens are lost ``` **Why bad:** Stream must be consumed via iteration, otherwise tokens are silently lost **See:** [examples/streaming.md](examples/streaming.md) for stream cancellation, controller access --- ### Pattern 4: Structured Output with JSON Schema Use `response_format: { type: "json_schema" }` with Zod-generated schemas. ```typescript import Together from "together-ai"; import { z } from "zod"; const client = new Together(); const EventSchema = z.object({ name: z.string(), date: z.string(), participants: z.array(z.string()), }); const jsonSchema = z.toJSONSchema(EventSchema); const completion = await client.chat.completions.create({ model: "Qwen/Qwen3.5-9B", messages: [ { role: "system", content: `Extract event details. Only answer in JSON. Follow this schema: ${JSON.stringify(jsonSchema)}`, }, { role: "user", content: "Alice and Bob meet next Tuesday for lunch." }, ], response_format: { type: "json_schema", json_schema: { name: "calendar_event", schema: jsonSchema }, }, }); const event = JSON.parse(completion.choices[0].message.content ?? "{}"); ``` **Why good:** Zod generates schema, schema included in both system prompt and `response_format`, named schema object ```typescript // BAD: Schema only in response_format, not in system prompt const completion = await client.chat.completions.create({ model: "Qwen/Qwen3.5-9B", messages: [{ role: "user", content: "Extract event details." }], response_format: { type: "json_schema", json_schema: { name: "event", schema: jsonSchema }, }, }); ``` **Why bad:** Model needs the schema in the system prompt AND `response_format` for reliable structured output -- omitting the prompt instruction degrades output quality **See:** [examples/structured-output.md](examples/structured-output.md) for regex mode, vision with JSON, complex schemas --- ### Pattern 5: Function Calling / Tool Use Define functions the model can call. Same `tools` parameter shape as OpenAI. ```typescript const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Weather in Paris?" }], tools: [ { type: "function", function: { name: "get_weather", description: "Get current weather for a location", parameters: { type: "object", properties: { location: { type: "string", description: "City name" }, }, required: ["location"], additionalProperties: false, }, strict: true, }, }, ], }); const toolCall = completion.choices[0].message.tool_calls?.[0]; if (toolCall) { const args = JSON.parse(toolCall.function.arguments); console.log(`Call ${toolCall.function.name} with:`, args); } ``` **Why good:** Standard OpenAI-compatible tool format, strict mode for reliable arguments, `additionalProperties: false` prevents hallucinated fields **See:** [examples/tools.md](examples/tools.md) for multi-step tool loops, `tool_choice`, parallel calls, supported models --- ### Pattern 6: Image Generation Generate images with FLUX and Stable Diffusion models. ```typescript const response = await client.images.generate({ model: "black-forest-labs/FLUX.1-schnell", prompt: "A serene mountain landscape at sunset with a lake reflection", steps: 4, }); console.log(response.data[0].url); ``` **Why good:** Simple API, model-specific parameters, URL response by default **See:** [examples/images.md](examples/images.md) for FLUX variants, base64, reference images, multiple variations --- ### Pattern 7: Embeddings Create embeddings for semantic search and RAG pipelines. ```typescript const EMBEDDING_MODEL = "BAAI/bge-large-en-v1.5"; const response = await client.embeddings.create({ model: EMBEDDING_MODEL, input: "TypeScript provides static type checking.", }); console.log(response.data[0].embedding); ``` **Why good:** Named model constant, simple single-input embedding, array response **See:** [examples/images.md](examples/images.md) for batch embeddings, semantic search with cosine similarity --- ### Pattern 8: Error Handling Always catch `Together.APIError` and its subclasses. ```typescript try { const completion = await client.chat.completions.create({ model: "meta-llama/Llama-3.3-70B-Instruct-Turbo", messages: [{ role: "user", content: "Hello" }], }); } catch (error) { if (error instanceof Together.APIError) { console.error(`API Error [${error.status}]: ${error.message}`); if (error instanceof Together.RateLimitError) { console.error("Rate limited -- SDK will auto-retry."); } if (error instanceof Together.AuthenticationError) { throw new Error("Invalid API key. Check TOGETHER_API_KEY."); } } else { throw error; // Re-throw non-API errors } } ``` **Why good:** Specific error types, re-throws unexpected errors, actionable error messages **See:** [examples/core.md](examples/core.md) for full production error handling, error type hierarchy </patterns> --- <performance> ## Performance Optimization ### Model Selection for Cost/Speed ``` Fast + cheap -> Llama 3.3 70B Turbo, Qwen3.5 9B Most capable -> DeepSeek V3.1, Qwen3.5 397B Complex reasoning -> DeepSeek R1 Function calling -> Llama 3.3 70B, Qwen3.5 9B, DeepSeek V3 Structured output (JSON) -> Qwen3.5 9B, Llama 3.3 70B Embeddings -> BAAI/bge-large-en-v1.5 (quality), UAE-Large-V1 Image generation (fast) -> FLUX.1 schnell (4 steps) Image generation (quality)-> FLUX.2 pro, FLUX.1.1 pro Vision / multimodal -> Qwen3-VL-8B-Instruct, Llama 3.2 Vision ``` ### Key Optimization Patterns - **Use Turbo variants** for chat models -- they are optimized for Together's infrastructure - **Set `temperature: 0`** for deterministic output when possible - **Batch embedding inputs** -- pass an array of strings to `client.embeddings.create()` instead of one at a time - **Use `steps: 4`** for FLUX.1 schnell images (higher steps have diminishing returns) - **Use streaming** for user-facing responses to reduce perceived latency </performance> --- <decision_framework> ## Decision Framework ### Which Model to Choose ``` What is your task? +-- General chat / instruction following -> Llama 3.3 70B Turbo (fast, cheap) +-- Most capable reasoning -> DeepSeek V3.1, Qwen3.5 397B +-- Complex math / chain-of-thought -> DeepSeek R1 +-- Function calling / tool use -> Llama 3.3 70B, Qwen3.5 9B +-- Structured JSON output -> Qwen3.5 9B (best JSON mode support) +-- Vision / image understanding -> Qwen3-VL-8B-Instruct +-- Code generation -> DeepSeek V3, Qwen Coder +-- Embeddings -> BAAI/bge-large-en-v1.5 (default) +-- Image generation (fast) -> FLUX.1 schnell +-- Image generation (quality) -> FLUX.2 pro, FLUX.1.1 pro ``` ### Together AI SDK vs OpenAI SDK ``` Do you ONLY use Together AI models? +-- YES -> Use together-ai package (purpose-built, full API coverage) +-- NO -> Do you also use OpenAI models? +-- YES -> Two options: | +-- Separate SDKs: together-ai for Together, openai for OpenAI | +-- OpenAI SDK only: Point baseURL to api.together.xyz/v1 +-- NO -> Use a provider-agnostic SDK ``` ### Streaming vs Non-Streaming ``` Is the response user-facing? +-- YES -> Use streaming (stream: true) +-- NO -> Use non-streaming +-- Background processing -> client.chat.completions.create() +-- Structured output -> Non-streaming with response_format ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Hardcoding `TOGETHER_API_KEY` instead of using environment variables (security breach risk) - Using bare `catch` blocks without checking `Together.APIError` (hides API errors) - Not consuming streams returned by `stream: true` (tokens are silently lost) - Using `JSON.parse()` on completion content without `response_format` (fragile, model may return non-JSON) - Omitting the schema from the system prompt when using `response_format: { type: "json_schema" }` (degrades output quality) **Medium Priority Issues:** - Not setting `maxRetries` / `timeout` for production deployments (default timeout is 1 minute) - Missing `system` role message (no system instruction means unpredictable behavior) - Using a model that does not support function calling with `tools` parameter (will silently fail or error) - Not checking if `tool_calls` is defined before accessing arguments - Using `width`/`height` with FLUX schnell/Kontext models (use `aspect_ratio` instead) **Common Mistakes:** - Using OpenAI model names (e.g., `gpt-4o`) with the Together AI SDK -- Together uses Hugging Face-style IDs like `meta-llama/Llama-3.3-70B-Instruct-Turbo` - Confusing `client.images.generate()` (Together) with `client.images.create()` (OpenAI) -- different method name - Forgetting to use `z.toJSONSchema()` (Zod v4) or `zodToJsonSchema()` (Zod v3) to convert schemas before passing to `response_format` - Using the `developer` role (OpenAI-specific) instead of `system` role with Together AI models - Passing `max_completion_tokens` instead of `max_tokens` -- Together uses `max_tokens` **Gotchas & Edge Cases:** - The SDK auto-retries on 429 (rate limit), 408, 409, and 5xx errors -- 2 retries by default. Disable with `maxRetries: 0`. - Model IDs are case-sensitive and follow the `org/model-name` format from Hugging Face. - Not all models support function calling. See [examples/tools.md](examples/tools.md) for the current supported list, or check the [official docs](https://docs.together.ai/docs/function-calling). - FLUX.1 schnell and Kontext models use `aspect_ratio` parameter; FLUX.1 Pro and FLUX.1.1 Pro use `width`/`height`. - Image generation returns URLs by default. Use `response_format: "base64"` for inline data. - The `response_format: { type: "json_schema" }` requires telling the model to "only answer in JSON" in the system prompt -- the schema alone is not sufficient. - Structured output uses `z.toJSONSchema()` (Zod v4) -- if using Zod v3, use `zodToJsonSchema()` from the `zod-to-json-schema` package. - Together AI's `client.images.generate()` is the method name, not `client.images.create()` like OpenAI. - Fine-tuning supports LoRA and full fine-tuning. File format is JSONL with `messages` array per line. - The OpenAI compatibility endpoint (`api.together.xyz/v1`) supports chat, embeddings, images, vision, function calling, and structured output -- but not fine-tuning or model management. </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST use the `together-ai` package (`import Together from "together-ai"`) -- NOT the OpenAI SDK -- unless explicitly building an OpenAI-compatible integration)** **(You MUST include the JSON schema in BOTH the `response_format` parameter AND the system prompt when using structured output -- the model needs both)** **(You MUST handle errors using `Together.APIError` and its subclasses -- never use bare catch blocks without error type checking)** **(You MUST never hardcode API keys -- always use environment variables via `process.env.TOGETHER_API_KEY`)** **Failure to follow these rules will produce insecure, unreliable, or incorrectly structured AI integrations.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.