Claude Skill

ai-infrastructure-huggingface-inference

Hugging Face Inference SDK patterns for TypeScript/Node.js — InferenceClient setup, chat completion, text generation, streaming, embeddings, image generation, audio transcription, translation, summarization, and Inference Endpoints

LLM Mart · 0 points · 0 views 8 listing impressions 0 install-command copies
Virus-scanned Reviewed automatically before listing.

Full trust report

Download agents-inc-skills-dist_plugins_ai-infrastructure-huggingface-inference_skills_ai-infrastructure-huggingface-inference-3a51ef5.zip · 14 KB
Part of agents-inc/skills — 130 skills

Install

skills CLI npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-infrastructure-huggingface-inference/skills/ai-infrastructure-huggingface-inference
Claude Code claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
Git git clone https://github.com/agents-inc/skills.git

The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.

Skill manifest

Hugging Face Inference Patterns

Quick Guide: Use @huggingface/inference (v4+) to access 200k+ ML models on the Hugging Face Hub. Use InferenceClient with chatCompletion() for OpenAI-compatible chat, textGeneration() for raw text completion, chatCompletionStream() for streaming, featureExtraction() for embeddings, textToImage() for image generation, and automaticSpeechRecognition() for audio transcription. Set provider to route through inference providers (Cerebras, Together, Groq, etc.) or use endpointUrl for dedicated Inference Endpoints.


<critical_requirements>

CRITICAL: Before Using This Skill

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST always pass an access token to InferenceClient -- never deploy without authentication)

(You MUST use chatCompletion() / chatCompletionStream() for conversational LLM tasks -- these follow the OpenAI-compatible message format)

(You MUST handle errors using InferenceClientError and its subclasses -- never use bare catch blocks without error type checking)

(You MUST specify a model parameter for every inference call -- there is no default model)

(You MUST never hardcode access tokens -- always use environment variables via process.env.HF_TOKEN)

</critical_requirements>


Auto-detection: Hugging Face, huggingface, @huggingface/inference, InferenceClient, HfInference, hf.chatCompletion, hf.textGeneration, hf.featureExtraction, hf.textToImage, hf.automaticSpeechRecognition, hf.translation, hf.summarization, hf.textToSpeech, chatCompletionStream, textGenerationStream, HF_TOKEN, inference provider, Inference Endpoints

When to use:

  • Accessing any of the 200k+ models hosted on the Hugging Face Hub
  • Running chat completion with open-source LLMs (Qwen, Mistral, Llama, etc.)
  • Generating embeddings with sentence-transformer models for semantic search
  • Generating images from text prompts (FLUX, Stable Diffusion)
  • Transcribing audio with automatic speech recognition models
  • Running translation, summarization, text classification, or NER tasks
  • Deploying models on dedicated Inference Endpoints for production use
  • Using third-party inference providers (Cerebras, Together, Groq, Replicate, etc.) through a unified API

Key patterns covered:

  • InferenceClient initialization and configuration
  • Chat Completion API (OpenAI-compatible messages format, streaming)
  • Text generation (raw completion, streaming)
  • Embeddings via feature extraction
  • Image generation (text-to-image)
  • Audio transcription (automatic speech recognition)
  • Translation, summarization, and text classification
  • Inference Endpoints (dedicated deployments)
  • Inference Providers (routing through third-party services)
  • Error handling with typed error classes

When NOT to use:

  • If you only use OpenAI models -- use the OpenAI SDK directly
  • If you need a provider-agnostic unified SDK with structured outputs and tool calling -- use a higher-level AI SDK
  • If you need to fine-tune or train models -- use the @huggingface/hub package or Python transformers

Examples Index




<decision_framework>

Decision Framework

Which Method to Use

What is your task?
+-- Conversational LLM (messages) -> chatCompletion() / chatCompletionStream()
+-- Raw text continuation -> textGeneration() / textGenerationStream()
+-- Embeddings for search/RAG -> featureExtraction()
+-- Image from text prompt -> textToImage()
+-- Speech to text -> automaticSpeechRecognition()
+-- Text to speech -> textToSpeech()
+-- Language translation -> translation()
+-- Summarize long text -> summarization()
+-- Classify text -> textClassification()
+-- Named entity recognition -> tokenClassification()
+-- Classify image -> imageClassification()
+-- Detect objects -> objectDetection()
+-- Caption an image -> imageToText()
+-- Answer questions from context -> questionAnswering()

Chat Completion vs Text Generation

Do you have a conversation with roles (system/user/assistant)?
+-- YES -> chatCompletion() / chatCompletionStream()
|   Uses OpenAI-compatible message format
|   Supports system messages, multi-turn
+-- NO -> Do you want to continue/complete a text prompt?
    +-- YES -> textGeneration() / textGenerationStream()
    |   Takes raw text input via 'inputs'
    +-- NO -> Use a task-specific method instead

Serverless vs Dedicated

What are your deployment needs?
+-- Prototyping / low volume -> Serverless Inference Providers (provider: "auto")
|   Free tier available, shared infrastructure, may have cold starts
+-- Production / high volume -> Inference Endpoints (endpointUrl)
|   Dedicated GPU, autoscaling, scale-to-zero, private infrastructure
+-- Local development -> Local endpoint (endpointUrl: "http://localhost:8080")
    Works with llama.cpp, Ollama, vLLM, TGI, LiteLLM

When to Use This SDK vs Others

Do you need access to 200k+ open-source models?
+-- YES -> Use @huggingface/inference
+-- NO -> Do you only use OpenAI models?
    +-- YES -> Not this skill's scope -- use the OpenAI SDK directly
    +-- NO -> Do you need structured outputs / tool calling?
        +-- YES -> Not this skill's scope -- use a higher-level AI SDK
        +-- NO -> @huggingface/inference works for most ML tasks

</decision_framework>


<red_flags>

RED FLAGS

High Priority Issues:

  • Hardcoding access tokens instead of using environment variables (security breach risk)
  • Using bare catch blocks without checking InferenceClientError types (hides API errors, loses debug info)
  • Omitting the model parameter -- always specify the model explicitly for predictable behavior (the SDK can pick a recommended model if omitted, but this is unreliable for production)
  • Not consuming chatCompletionStream() / textGenerationStream() generators (tokens are silently lost)
  • Using textGeneration() for conversational tasks instead of chatCompletion() (wrong API shape)

Medium Priority Issues:

  • Not specifying provider when a specific provider is needed (default "auto" picks based on your HF settings, which may not be optimal)
  • Not checking chunk.choices length before accessing chunk.choices[0] in streaming (may throw on empty chunks)
  • Using request() / streamingRequest() directly -- these are deprecated, use task-specific methods
  • Ignoring max_tokens / max_new_tokens limits (output may be truncated or excessively long)
  • Not handling model loading time for serverless inference (cold models return 503, then load)

Common Mistakes:

  • Confusing chatCompletion() parameters with textGeneration() parameters -- chat uses messages + max_tokens, text generation uses inputs + parameters.max_new_tokens
  • Using inputs parameter with chatCompletion() -- it uses messages, not inputs
  • Using messages parameter with textGeneration() -- it uses inputs, not messages
  • Forgetting that textToImage() returns a Blob, not a URL or Buffer
  • Treating featureExtraction() output as always a flat array -- shape depends on the model (can be nested arrays for batch inputs)
  • Not passing data (binary) for audio/image tasks, or passing a string path instead of the actual file buffer

Gotchas & Edge Cases:

  • The provider: "auto" default selects providers based on your HF account settings at hf.co/settings/inference-providers -- not by availability or speed. Set an explicit provider for predictable routing.
  • Serverless models may need time to load (cold start). First requests to a cold model may return 503 errors while the model warms up. The SDK handles retries, but initial requests can be slow.
  • When using Inference Endpoints with endpointUrl, the model parameter is often ignored because the endpoint serves a specific model.
  • chatCompletion() is OpenAI-API compatible -- it works with any OpenAI-compatible endpoint, not just Hugging Face.
  • HfInference is still exported for backward compatibility but InferenceClient is the current class name.
  • Third-party provider API keys can be passed as the accessToken -- when authenticated with a non-HF key, requests go directly to the provider instead of through HF's routing layer.
  • Tree-shakeable imports (import { textGeneration } from "@huggingface/inference") require passing accessToken as a parameter instead of constructor.
  • textToImage() supports multiple output types: blob (default), url, dataUrl, or json via the options outputType parameter.
  • Translation requires parameters.src_lang and parameters.tgt_lang for many-to-many models like mbart-large-50-many-to-many-mmt.

</red_flags>


<critical_reminders>

CRITICAL REMINDERS

All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering, import type, named constants)

(You MUST always pass an access token to InferenceClient -- never deploy without authentication)

(You MUST use chatCompletion() / chatCompletionStream() for conversational LLM tasks -- these follow the OpenAI-compatible message format)

(You MUST handle errors using InferenceClientError and its subclasses -- never use bare catch blocks without error type checking)

(You MUST specify a model parameter for every inference call -- there is no default model)

(You MUST never hardcode access tokens -- always use environment variables via process.env.HF_TOKEN)

Failure to follow these rules will produce insecure, unreliable, or silently failing AI integrations.

</critical_reminders>

Files (skills)
  • examples
    • core.md 10.3 KB
      # Hugging Face Inference -- Setup, Chat & Text Generation Examples
      
      > Client initialization, chat completion, text generation, streaming, providers, endpoints, and error handling. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Related examples:**
      
      - [tasks.md](tasks.md) -- Embeddings, image generation, audio, translation, summarization
      
      ---
      
      ## Basic Client Setup
      
      ```typescript
      // lib/hf-client.ts
      import { InferenceClient } from "@huggingface/inference";
      
      // Requires HF_TOKEN environment variable
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      export { client };
      ```
      
      ---
      
      ## With Specific Provider
      
      ```typescript
      // lib/hf-client.ts
      import { InferenceClient } from "@huggingface/inference";
      
      // HF token routes requests through HF's proxy to the provider
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      // Provider specified per-request
      const response = await client.chatCompletion({
        model: "Qwen/Qwen3-32B",
        provider: "cerebras", // Route to Cerebras infrastructure
        messages: [{ role: "user", content: "Hello!" }],
      });
      ```
      
      ---
      
      ## With Dedicated Inference Endpoint
      
      ```typescript
      // lib/hf-endpoint.ts
      import { InferenceClient } from "@huggingface/inference";
      
      // Option 1: endpointUrl in constructor
      const ENDPOINT_URL =
        "https://j3z5luu0ooo76jnl.us-east-1.aws.endpoints.huggingface.cloud/v1/";
      
      const client = new InferenceClient(process.env.HF_TOKEN, {
        endpointUrl: ENDPOINT_URL,
      });
      
      // Model parameter is ignored -- endpoint serves a specific model
      const response = await client.chatCompletion({
        messages: [{ role: "user", content: "What is the capital of France?" }],
      });
      
      console.log(response.choices[0].message.content);
      
      export { client };
      ```
      
      ```typescript
      // Option 2: .endpoint() helper
      import { InferenceClient } from "@huggingface/inference";
      
      const ENDPOINT_URL =
        "https://j3z5luu0ooo76jnl.us-east-1.aws.endpoints.huggingface.cloud/v1/";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      const endpointClient = client.endpoint(ENDPOINT_URL);
      
      const response = await endpointClient.chatCompletion({
        messages: [{ role: "user", content: "Hello!" }],
      });
      ```
      
      ---
      
      ## With Local Endpoint (Ollama, llama.cpp, vLLM, TGI, LiteLLM)
      
      ```typescript
      // lib/hf-local.ts
      import { InferenceClient } from "@huggingface/inference";
      
      // No token needed for local endpoints
      const client = new InferenceClient(undefined, {
        endpointUrl: "http://localhost:8080",
      });
      
      const response = await client.chatCompletion({
        messages: [{ role: "user", content: "What is the capital of France?" }],
      });
      
      console.log(response.choices[0].message.content);
      ```
      
      ---
      
      ## Disable Endpoint Retry-on-Error
      
      ```typescript
      // By default, calls wait for model to load (scale-to-zero endpoints)
      // Disable to handle 503 errors yourself
      const response = await client.chatCompletion(
        {
          messages: [{ role: "user", content: "Hello" }],
        },
        {
          retry_on_error: false,
        },
      );
      ```
      
      ---
      
      ## Chat Completion -- Basic
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      const MAX_TOKENS = 512;
      const TEMPERATURE = 0.1;
      
      async function chat(userMessage: string): Promise<string> {
        const response = await client.chatCompletion({
          model: "Qwen/Qwen3-32B",
          provider: "cerebras",
          messages: [
            { role: "system", content: "You are a helpful assistant. Be concise." },
            { role: "user", content: userMessage },
          ],
          max_tokens: MAX_TOKENS,
          temperature: TEMPERATURE,
        });
      
        const content = response.choices[0].message.content;
        if (!content) {
          throw new Error("No content in response");
        }
      
        return content;
      }
      
      const answer = await chat("What is TypeScript in one sentence?");
      console.log(answer);
      ```
      
      ---
      
      ## Chat Completion -- Multi-Turn
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      const MAX_TOKENS = 512;
      
      interface Message {
        role: "system" | "user" | "assistant";
        content: string;
      }
      
      const messages: Message[] = [
        { role: "system", content: "You are a TypeScript expert." },
        { role: "user", content: "What is a union type?" },
      ];
      
      const response = await client.chatCompletion({
        model: "Qwen/Qwen3-32B",
        provider: "cerebras",
        messages,
        max_tokens: MAX_TOKENS,
      });
      
      // Append assistant response for next turn
      const assistantContent = response.choices[0].message.content ?? "";
      messages.push({ role: "assistant", content: assistantContent });
      messages.push({ role: "user", content: "Give me a real-world example." });
      
      const followUp = await client.chatCompletion({
        model: "Qwen/Qwen3-32B",
        provider: "cerebras",
        messages,
        max_tokens: MAX_TOKENS,
      });
      ```
      
      ---
      
      ## Chat Completion -- Streaming
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      const MAX_TOKENS = 512;
      
      async function streamChat(prompt: string): Promise<string> {
        let fullResponse = "";
      
        for await (const chunk of client.chatCompletionStream({
          model: "Qwen/Qwen3-32B",
          provider: "cerebras",
          messages: [
            { role: "system", content: "You are a helpful assistant." },
            { role: "user", content: prompt },
          ],
          max_tokens: MAX_TOKENS,
        })) {
          if (chunk.choices && chunk.choices.length > 0) {
            const content = chunk.choices[0].delta.content;
            if (content) {
              process.stdout.write(content);
              fullResponse += content;
            }
          }
        }
      
        console.log(); // newline
        return fullResponse;
      }
      
      const result = await streamChat("Explain promises in JavaScript.");
      console.log("Response length:", result.length);
      ```
      
      ---
      
      ## Text Generation -- Non-Streaming
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      const MAX_NEW_TOKENS = 250;
      
      const result = await client.textGeneration({
        model: "mistralai/Mixtral-8x7B-v0.1",
        provider: "together",
        inputs: "The key benefits of TypeScript are",
        parameters: { max_new_tokens: MAX_NEW_TOKENS },
      });
      
      console.log(result.generated_text);
      ```
      
      ---
      
      ## Text Generation -- Streaming
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      const MAX_NEW_TOKENS = 250;
      
      for await (const output of client.textGenerationStream({
        model: "mistralai/Mixtral-8x7B-v0.1",
        provider: "together",
        inputs: "The key benefits of TypeScript are",
        parameters: { max_new_tokens: MAX_NEW_TOKENS },
      })) {
        if (output.token.text) {
          process.stdout.write(output.token.text);
        }
      }
      console.log();
      ```
      
      ---
      
      ## Tree-Shakeable Imports
      
      ```typescript
      // Import individual functions for smaller bundles
      import { textGeneration, chatCompletion } from "@huggingface/inference";
      
      const MAX_NEW_TOKENS = 250;
      
      // Each function requires accessToken as a parameter
      const result = await textGeneration({
        accessToken: process.env.HF_TOKEN,
        model: "mistralai/Mixtral-8x7B-v0.1",
        provider: "together",
        inputs: "The key benefits of TypeScript are",
        parameters: { max_new_tokens: MAX_NEW_TOKENS },
      });
      ```
      
      ---
      
      ## Third-Party Provider API Key (Direct)
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      // Using a provider's own API key bypasses HF routing -- goes directly to provider
      const client = new InferenceClient(process.env.MISTRAL_API_KEY, {
        endpointUrl: "https://api.mistral.ai",
      });
      
      const response = await client.chatCompletion({
        model: "mistral-tiny",
        messages: [{ role: "user", content: "Hello!" }],
      });
      ```
      
      ---
      
      ## Error Handling -- Production Pattern
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      import {
        InferenceClientError,
        InferenceClientInputError,
        InferenceClientProviderApiError,
        InferenceClientProviderOutputError,
        InferenceClientHubApiError,
      } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      const MAX_TOKENS = 512;
      
      async function safeChat(prompt: string): Promise<string | null> {
        try {
          const response = await client.chatCompletion({
            model: "Qwen/Qwen3-32B",
            provider: "cerebras",
            messages: [
              { role: "system", content: "You are a helpful assistant." },
              { role: "user", content: prompt },
            ],
            max_tokens: MAX_TOKENS,
          });
      
          return response.choices[0].message.content;
        } catch (error) {
          if (error instanceof InferenceClientProviderApiError) {
            // API-level errors from the inference provider (rate limits, auth, server errors)
            console.error("Provider API error:", error.message);
            console.error("Request details:", error.request);
            console.error("Response details:", error.response);
            return null;
          }
      
          if (error instanceof InferenceClientHubApiError) {
            // Errors from the HF Hub API (model not found, repo issues)
            console.error("Hub API error:", error.message);
            return null;
          }
      
          if (error instanceof InferenceClientProviderOutputError) {
            // Provider returned malformed response
            console.error("Malformed provider response:", error.message);
            return null;
          }
      
          if (error instanceof InferenceClientInputError) {
            // Invalid input parameters
            console.error("Invalid input:", error.message);
            return null;
          }
      
          // Catch-all for any @huggingface/inference error
          if (error instanceof InferenceClientError) {
            console.error("Inference error:", error.message);
            return null;
          }
      
          // Unknown errors should be re-thrown
          throw error;
        }
      }
      
      const result = await safeChat("Hello!");
      if (result) {
        console.log(result);
      } else {
        console.error("Failed to get response");
      }
      ```
      
      ---
      
      ## Error Type Reference
      
      ```typescript
      // Error class hierarchy:
      // InferenceClientError (base)
      //   +-- InferenceClientInputError         (invalid input parameters)
      //   +-- InferenceClientProviderApiError    (provider API errors: rate limits, auth, 5xx)
      //   |     .request  -- Request details (URL, method, headers)
      //   |     .response -- Response details (status code, body)
      //   +-- InferenceClientHubApiError         (HF Hub API errors: model not found)
      //   |     .request  -- Request details
      //   |     .response -- Response details
      //   +-- InferenceClientProviderOutputError (malformed provider response)
      ```
      
      ---
      
      _For task-specific examples (embeddings, images, audio), see [tasks.md](tasks.md). For API reference tables, see [reference.md](../reference.md)._
      
    • tasks.md 10.9 KB
      # Hugging Face Inference -- Task Examples (Embeddings, Vision, Audio, NLP)
      
      > Task-specific examples: feature extraction, image generation, speech recognition, translation, summarization, classification, and more. See [SKILL.md](../SKILL.md) for core patterns.
      
      **Prerequisites**: Understand client setup and chat completion from [core.md](core.md) first.
      
      **Related examples:**
      
      - [core.md](core.md) -- Client setup, chat completion, text generation, streaming
      
      ---
      
      ## Feature Extraction (Embeddings)
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      // Single input -- returns number[]
      const embedding = await client.featureExtraction({
        model: "sentence-transformers/all-MiniLM-L6-v2",
        inputs: "That is a happy person",
      });
      ```
      
      ---
      
      ## Batch Embeddings with Cosine Similarity
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      function cosineSimilarity(a: number[], b: number[]): number {
        let dotProduct = 0;
        let normA = 0;
        let normB = 0;
        for (let i = 0; i < a.length; i++) {
          dotProduct += a[i] * b[i];
          normA += a[i] * a[i];
          normB += b[i] * b[i];
        }
        return dotProduct / (Math.sqrt(normA) * Math.sqrt(normB));
      }
      
      const texts = [
        "TypeScript adds static types to JavaScript",
        "JavaScript is a dynamic programming language",
        "The weather is nice today",
      ];
      
      // featureExtraction accepts a single string input
      // For batch, call individually or use a model that supports array inputs
      const embeddings = await Promise.all(
        texts.map((text) =>
          client.featureExtraction({
            model: "sentence-transformers/all-MiniLM-L6-v2",
            inputs: text,
          }),
        ),
      );
      
      // Compare first text to all others
      const SIMILARITY_THRESHOLD = 0.5;
      for (let i = 1; i < texts.length; i++) {
        const similarity = cosineSimilarity(
          embeddings[0] as number[],
          embeddings[i] as number[],
        );
        const isRelated = similarity > SIMILARITY_THRESHOLD;
        console.log(
          `"${texts[0]}" vs "${texts[i]}": ${similarity.toFixed(4)} (${isRelated ? "related" : "unrelated"})`,
        );
      }
      ```
      
      ---
      
      ## Text-to-Image
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      import { writeFileSync } from "node:fs";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      // Default output is Blob
      const imageBlob = await client.textToImage({
        model: "black-forest-labs/FLUX.1-dev",
        inputs: "a serene mountain landscape at sunset, oil painting style",
        provider: "replicate",
      });
      
      // Write Blob to file
      const buffer = Buffer.from(await imageBlob.arrayBuffer());
      writeFileSync("output/landscape.png", buffer);
      ```
      
      ---
      
      ## Text-to-Image with Output Options
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      // Get as URL string
      const imageUrl = await client.textToImage(
        {
          model: "black-forest-labs/FLUX.1-dev",
          inputs: "a cat wearing a hat",
        },
        { outputType: "url" },
      );
      console.log("Image URL:", imageUrl);
      
      // Get as data URL (base64 embedded)
      const dataUrl = await client.textToImage(
        {
          model: "black-forest-labs/FLUX.1-dev",
          inputs: "a cat wearing a hat",
        },
        { outputType: "dataUrl" },
      );
      // Use dataUrl directly in <img src="...">
      ```
      
      ---
      
      ## Image Classification
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      import { readFileSync } from "node:fs";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const results = await client.imageClassification({
        model: "google/vit-base-patch16-224",
        data: readFileSync("images/photo.png"),
      });
      
      // Results: Array<{ label: string, score: number }>
      for (const result of results) {
        console.log(`${result.label}: ${(result.score * 100).toFixed(1)}%`);
      }
      ```
      
      ---
      
      ## Image to Text (Captioning)
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      import { readFileSync } from "node:fs";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const result = await client.imageToText({
        model: "nlpconnect/vit-gpt2-image-captioning",
        data: readFileSync("images/photo.png"),
      });
      
      console.log("Caption:", result.generated_text);
      ```
      
      ---
      
      ## Object Detection
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      import { readFileSync } from "node:fs";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const detections = await client.objectDetection({
        model: "facebook/detr-resnet-50",
        data: readFileSync("images/street.png"),
      });
      
      // Results: Array<{ label: string, score: number, box: { xmin, ymin, xmax, ymax } }>
      for (const detection of detections) {
        console.log(
          `${detection.label} (${(detection.score * 100).toFixed(1)}%) at [${detection.box.xmin}, ${detection.box.ymin}]`,
        );
      }
      ```
      
      ---
      
      ## Automatic Speech Recognition (Transcription)
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      import { readFileSync } from "node:fs";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const result = await client.automaticSpeechRecognition({
        model: "facebook/wav2vec2-large-960h-lv60-self",
        data: readFileSync("audio/recording.flac"),
      });
      
      console.log("Transcript:", result.text);
      ```
      
      ---
      
      ## Audio Classification
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      import { readFileSync } from "node:fs";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const results = await client.audioClassification({
        model: "superb/hubert-large-superb-er",
        data: readFileSync("audio/sample.flac"),
      });
      
      // Results: Array<{ label: string, score: number }>
      for (const result of results) {
        console.log(`${result.label}: ${(result.score * 100).toFixed(1)}%`);
      }
      ```
      
      ---
      
      ## Text to Speech
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      import { writeFileSync } from "node:fs";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const audioBlob = await client.textToSpeech({
        model: "espnet/kan-bayashi_ljspeech_vits",
        inputs: "Hello, welcome to the application!",
      });
      
      const buffer = Buffer.from(await audioBlob.arrayBuffer());
      writeFileSync("output/speech.wav", buffer);
      ```
      
      ---
      
      ## Translation
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      // Simple translation (model determines source/target)
      const result = await client.translation({
        model: "t5-base",
        inputs: "My name is Wolfgang and I live in Berlin",
      });
      
      console.log("Translation:", result.translation_text);
      ```
      
      ```typescript
      // Many-to-many model with explicit language codes
      const result = await client.translation({
        model: "facebook/mbart-large-50-many-to-many-mmt",
        inputs: "My name is Wolfgang and I live in Berlin",
        parameters: {
          src_lang: "en_XX",
          tgt_lang: "fr_XX",
        },
      });
      
      console.log("French:", result.translation_text);
      ```
      
      ---
      
      ## Summarization
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      const MAX_LENGTH = 100;
      
      const result = await client.summarization({
        model: "facebook/bart-large-cnn",
        inputs:
          "The tower is 324 metres (1,063 ft) tall, about the same height as an 81-storey building, " +
          "and the tallest structure in Paris. Its base is square, measuring 125 metres (410 ft) on " +
          "each side. During its construction, the Eiffel Tower surpassed the Washington Monument to " +
          "become the tallest man-made structure in the world, a title it held for 41 years until the " +
          "Chrysler Building in New York City was finished in 1930.",
        parameters: { max_length: MAX_LENGTH },
      });
      
      console.log("Summary:", result.summary_text);
      ```
      
      ---
      
      ## Text Classification (Sentiment Analysis)
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const results = await client.textClassification({
        model: "distilbert-base-uncased-finetuned-sst-2-english",
        inputs: "I love this product! It works perfectly.",
      });
      
      // Results: Array<{ label: string, score: number }>
      for (const result of results) {
        console.log(`${result.label}: ${(result.score * 100).toFixed(1)}%`);
      }
      ```
      
      ---
      
      ## Token Classification (Named Entity Recognition)
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const entities = await client.tokenClassification({
        model: "dbmdz/bert-large-cased-finetuned-conll03-english",
        inputs: "My name is Sarah Jessica Parker but you can call me Jessica",
      });
      
      // Results: Array<{ entity_group: string, word: string, score: number, start: number, end: number }>
      for (const entity of entities) {
        console.log(
          `${entity.word} -> ${entity.entity_group} (${(entity.score * 100).toFixed(1)}%)`,
        );
      }
      ```
      
      ---
      
      ## Zero-Shot Classification
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const result = await client.zeroShotClassification({
        model: "facebook/bart-large-mnli",
        inputs: [
          "Hi, I recently bought a device from your company but it is not working as advertised and I would like to get reimbursed!",
        ],
        parameters: { candidate_labels: ["refund", "legal", "faq"] },
      });
      
      // Result includes labels sorted by score
      console.log(result);
      ```
      
      ---
      
      ## Question Answering (Extractive)
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const result = await client.questionAnswering({
        model: "deepset/roberta-base-squad2",
        inputs: {
          question: "What is the capital of France?",
          context:
            "The capital of France is Paris. It is located in the north of the country.",
        },
      });
      
      console.log(`Answer: ${result.answer} (score: ${result.score.toFixed(4)})`);
      ```
      
      ---
      
      ## Sentence Similarity
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const result = await client.sentenceSimilarity({
        model: "sentence-transformers/paraphrase-xlm-r-multilingual-v1",
        inputs: {
          source_sentence: "That is a happy person",
          sentences: [
            "That is a happy dog",
            "That is a very happy person",
            "Today is a sunny day",
          ],
        },
      });
      
      // Returns: number[] (similarity scores for each sentence)
      console.log("Similarity scores:", result);
      ```
      
      ---
      
      ## Visual Question Answering
      
      ```typescript
      import { InferenceClient } from "@huggingface/inference";
      
      const client = new InferenceClient(process.env.HF_TOKEN);
      
      const result = await client.visualQuestionAnswering({
        model: "dandelin/vilt-b32-finetuned-vqa",
        inputs: {
          question: "How many cats are in the image?",
          image: await (await fetch("https://example.com/cats.jpg")).blob(),
        },
      });
      
      console.log(`Answer: ${result.answer} (score: ${result.score.toFixed(4)})`);
      ```
      
      ---
      
      _For client setup and chat patterns, see [core.md](core.md). For API reference tables, see [reference.md](../reference.md)._
      
  • reference.md 10.9 KB
    # Hugging Face Inference Quick Reference
    
    > Method signatures, error types, provider list, and model recommendations. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples.
    
    ---
    
    ## Package Installation
    
    ```bash
    npm install @huggingface/inference
    ```
    
    ---
    
    ## Client Initialization
    
    ```typescript
    import { InferenceClient } from "@huggingface/inference";
    
    const client = new InferenceClient(accessToken, {
      endpointUrl?: string,  // Custom endpoint URL (Inference Endpoints or local)
    });
    ```
    
    ### Environment Variables
    
    | Variable   | Purpose                              |
    | ---------- | ------------------------------------ |
    | `HF_TOKEN` | Hugging Face access token (required) |
    
    ---
    
    ## API Methods Reference
    
    ### NLP -- Chat & Text Generation
    
    ```typescript
    // Chat Completion (OpenAI-compatible)
    const response = await client.chatCompletion({
      model: string,            // Required: model ID on the Hub
      messages: Message[],      // Required: { role, content }[]
      max_tokens?: number,      // Max output tokens
      temperature?: number,     // 0-2 (default: 1)
      top_p?: number,           // Nucleus sampling
      provider?: string,        // Inference provider (default: "auto")
    });
    // Returns: { choices: [{ message: { role, content } }] }
    
    // Chat Completion Streaming
    const stream = client.chatCompletionStream({
      model: string,
      messages: Message[],
      max_tokens?: number,
      temperature?: number,
      provider?: string,
    });
    // Returns: AsyncGenerator<{ choices: [{ delta: { content } }] }>
    
    // Text Generation
    const result = await client.textGeneration({
      model: string,             // Required
      inputs: string,            // Required: prompt text
      parameters?: {
        max_new_tokens?: number, // Max tokens to generate
        temperature?: number,
        top_p?: number,
        repetition_penalty?: number,
      },
      provider?: string,
    });
    // Returns: { generated_text: string }
    
    // Text Generation Streaming
    const stream = client.textGenerationStream({
      model: string,
      inputs: string,
      parameters?: { ... },
      provider?: string,
    });
    // Returns: AsyncGenerator<{ token: { text, id, logprob }, generated_text?: string }>
    ```
    
    ### NLP -- Analysis Tasks
    
    ```typescript
    // Feature Extraction (Embeddings)
    await client.featureExtraction({
      model: string,    // Required
      inputs: string,   // Required: text to embed
    });
    // Returns: number[] | number[][]
    
    // Summarization
    await client.summarization({
      model: string,
      inputs: string,
      parameters?: { max_length?: number, min_length?: number },
    });
    // Returns: { summary_text: string }
    
    // Translation
    await client.translation({
      model: string,
      inputs: string,
      parameters?: { src_lang?: string, tgt_lang?: string },
    });
    // Returns: { translation_text: string }
    
    // Text Classification
    await client.textClassification({
      model: string,
      inputs: string,
    });
    // Returns: Array<{ label: string, score: number }>
    
    // Token Classification (NER)
    await client.tokenClassification({
      model: string,
      inputs: string,
    });
    // Returns: Array<{ entity_group: string, word: string, score: number, start: number, end: number }>
    
    // Zero-Shot Classification
    await client.zeroShotClassification({
      model: string,
      inputs: string[],
      parameters: { candidate_labels: string[] },
    });
    
    // Question Answering
    await client.questionAnswering({
      model: string,
      inputs: { question: string, context: string },
    });
    // Returns: { answer: string, score: number, start: number, end: number }
    
    // Fill Mask
    await client.fillMask({
      model: string,
      inputs: string,   // Must contain [MASK] token
    });
    
    // Sentence Similarity
    await client.sentenceSimilarity({
      model: string,
      inputs: { source_sentence: string, sentences: string[] },
    });
    // Returns: number[]
    ```
    
    ### Audio
    
    ```typescript
    // Automatic Speech Recognition
    await client.automaticSpeechRecognition({
      model: string,
      data: Blob | ArrayBuffer, // Audio file data
    });
    // Returns: { text: string }
    
    // Audio Classification
    await client.audioClassification({
      model: string,
      data: Blob | ArrayBuffer,
    });
    // Returns: Array<{ label: string, score: number }>
    
    // Text to Speech
    await client.textToSpeech({
      model: string,
      inputs: string,
    });
    // Returns: Blob (audio)
    
    // Audio to Audio
    await client.audioToAudio({
      model: string,
      data: Blob | ArrayBuffer,
    });
    // Returns: Array<{ blob: Blob, label: string, "content-type": string }>
    ```
    
    ### Computer Vision
    
    ```typescript
    // Text to Image
    await client.textToImage({
      model: string,
      inputs: string,        // Text prompt
      provider?: string,
    }, {
      outputType?: "blob" | "url" | "dataUrl" | "json",  // Default: "blob"
    });
    // Returns: Blob | string (depends on outputType)
    
    // Image Classification
    await client.imageClassification({
      model: string,
      data: Blob | ArrayBuffer,
    });
    // Returns: Array<{ label: string, score: number }>
    
    // Object Detection
    await client.objectDetection({
      model: string,
      data: Blob | ArrayBuffer,
    });
    // Returns: Array<{ label: string, score: number, box: { xmin, ymin, xmax, ymax } }>
    
    // Image to Text (Captioning)
    await client.imageToText({
      model: string,
      data: Blob | ArrayBuffer,
    });
    // Returns: { generated_text: string }
    
    // Image Segmentation
    await client.imageSegmentation({
      model: string,
      data: Blob | ArrayBuffer,
    });
    // Returns: Array<{ label: string, score: number, mask: string }>
    
    // Image to Image
    await client.imageToImage({
      inputs: Blob,
      parameters?: { prompt?: string },
      model: string,
    });
    // Returns: Blob
    
    // Zero-Shot Image Classification
    await client.zeroShotImageClassification({
      model: string,
      inputs: { image: Blob },
      parameters: { candidate_labels: string[] },
    });
    ```
    
    ### Multimodal
    
    ```typescript
    // Visual Question Answering
    await client.visualQuestionAnswering({
      model: string,
      inputs: { question: string, image: Blob },
    });
    // Returns: { answer: string, score: number }
    
    // Document Question Answering
    await client.documentQuestionAnswering({
      model: string,
      inputs: { question: string, image: Blob },
    });
    ```
    
    ---
    
    ## Recommended Models by Task
    
    | Task                       | Model                                              | Notes                    |
    | -------------------------- | -------------------------------------------------- | ------------------------ |
    | Chat Completion            | `Qwen/Qwen3-32B`                                   | Strong open-source LLM   |
    | Chat Completion            | `mistralai/Mixtral-8x7B-v0.1`                      | Mixture of experts       |
    | Text Generation            | `mistralai/Mixtral-8x7B-v0.1`                      | Text continuation        |
    | Embeddings                 | `sentence-transformers/all-MiniLM-L6-v2`           | Fast, good quality       |
    | Multilingual Embeddings    | `intfloat/multilingual-e5-large`                   | Cross-lingual            |
    | Text-to-Image              | `black-forest-labs/FLUX.1-dev`                     | High quality image gen   |
    | Image Classification       | `google/vit-base-patch16-224`                      | General purpose          |
    | Object Detection           | `facebook/detr-resnet-50`                          | General purpose          |
    | Speech Recognition         | `facebook/wav2vec2-large-960h-lv60-self`           | English                  |
    | Audio Classification       | `superb/hubert-large-superb-er`                    | Emotion recognition      |
    | Text-to-Speech             | `espnet/kan-bayashi_ljspeech_vits`                 | English TTS              |
    | Translation                | `t5-base`                                          | English focused          |
    | Translation (multilingual) | `facebook/mbart-large-50-many-to-many-mmt`         | 50 languages             |
    | Summarization              | `facebook/bart-large-cnn`                          | News summarization       |
    | Sentiment Analysis         | `distilbert-base-uncased-finetuned-sst-2-english`  | Binary sentiment         |
    | NER                        | `dbmdz/bert-large-cased-finetuned-conll03-english` | Named entities           |
    | Zero-Shot Classification   | `facebook/bart-large-mnli`                         | No training needed       |
    | Question Answering         | `deepset/roberta-base-squad2`                      | Extractive QA            |
    | Image Captioning           | `nlpconnect/vit-gpt2-image-captioning`             | General captioning       |
    | Fill Mask                  | `bert-base-uncased`                                | Masked language modeling |
    | Zero-Shot Image Class.     | `openai/clip-vit-large-patch14-336`                | No training needed       |
    
    ---
    
    ## Error Types
    
    | Error Class                          | Cause                                        | Has `.request` / `.response`? |
    | ------------------------------------ | -------------------------------------------- | ----------------------------- |
    | `InferenceClientError`               | Base class for all errors                    | No                            |
    | `InferenceClientInputError`          | Invalid input parameters                     | No                            |
    | `InferenceClientProviderApiError`    | Provider API errors (rate limits, auth, 5xx) | Yes                           |
    | `InferenceClientHubApiError`         | HF Hub API errors (model not found)          | Yes                           |
    | `InferenceClientProviderOutputError` | Malformed provider response                  | No                            |
    
    All errors extend the base `Error` class via `InferenceClientError`.
    
    ---
    
    ## Supported Inference Providers
    
    | Provider        | Key Tasks Supported              |
    | --------------- | -------------------------------- |
    | Baseten         | Chat, text generation            |
    | Blackforestlabs | Image generation                 |
    | Cerebras        | Chat completion, text generation |
    | Clarifai        | Chat, text, image generation     |
    | Cohere          | Chat, text, embeddings           |
    | DeepInfra       | Chat, text, embeddings           |
    | Fal.ai          | Image, video generation          |
    | Featherless AI  | Chat, text generation            |
    | Fireworks AI    | Chat, text generation            |
    | Groq            | Chat completion (fast inference) |
    | HF Inference    | All tasks (Hugging Face's own)   |
    | Hyperbolic      | Chat, text generation            |
    | Nebius          | Chat, text generation            |
    | Novita          | Chat, image generation           |
    | Nscale          | Chat, text generation            |
    | NVIDIA          | Chat, text, embeddings           |
    | OVHcloud        | Chat, text generation            |
    | Public AI       | Chat, text generation            |
    | Replicate       | Image generation, audio          |
    | Sambanova       | Chat, text generation            |
    | Scaleway        | Chat, text generation            |
    | Together        | Chat, text, image generation     |
    | Wavespeed.ai    | Image, video generation          |
    | Z.ai            | Chat, text generation            |
    
    Set `provider: "auto"` (default) to use your account's preferred provider order.
    
    ---
    
    See [examples/core.md](examples/core.md) for tree-shakeable imports and endpoint configuration patterns.
    
  • SKILL.md 18 KB
    ---
    name: ai-infrastructure-huggingface-inference
    description: Hugging Face Inference SDK patterns for TypeScript/Node.js — InferenceClient setup, chat completion, text generation, streaming, embeddings, image generation, audio transcription, translation, summarization, and Inference Endpoints
    ---
    
    # Hugging Face Inference Patterns
    
    > **Quick Guide:** Use `@huggingface/inference` (v4+) to access 200k+ ML models on the Hugging Face Hub. Use `InferenceClient` with `chatCompletion()` for OpenAI-compatible chat, `textGeneration()` for raw text completion, `chatCompletionStream()` for streaming, `featureExtraction()` for embeddings, `textToImage()` for image generation, and `automaticSpeechRecognition()` for audio transcription. Set `provider` to route through inference providers (Cerebras, Together, Groq, etc.) or use `endpointUrl` for dedicated Inference Endpoints.
    
    ---
    
    <critical_requirements>
    
    ## CRITICAL: Before Using This Skill
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST always pass an access token to `InferenceClient` -- never deploy without authentication)**
    
    **(You MUST use `chatCompletion()` / `chatCompletionStream()` for conversational LLM tasks -- these follow the OpenAI-compatible message format)**
    
    **(You MUST handle errors using `InferenceClientError` and its subclasses -- never use bare catch blocks without error type checking)**
    
    **(You MUST specify a `model` parameter for every inference call -- there is no default model)**
    
    **(You MUST never hardcode access tokens -- always use environment variables via `process.env.HF_TOKEN`)**
    
    </critical_requirements>
    
    ---
    
    **Auto-detection:** Hugging Face, huggingface, @huggingface/inference, InferenceClient, HfInference, hf.chatCompletion, hf.textGeneration, hf.featureExtraction, hf.textToImage, hf.automaticSpeechRecognition, hf.translation, hf.summarization, hf.textToSpeech, chatCompletionStream, textGenerationStream, HF_TOKEN, inference provider, Inference Endpoints
    
    **When to use:**
    
    - Accessing any of the 200k+ models hosted on the Hugging Face Hub
    - Running chat completion with open-source LLMs (Qwen, Mistral, Llama, etc.)
    - Generating embeddings with sentence-transformer models for semantic search
    - Generating images from text prompts (FLUX, Stable Diffusion)
    - Transcribing audio with automatic speech recognition models
    - Running translation, summarization, text classification, or NER tasks
    - Deploying models on dedicated Inference Endpoints for production use
    - Using third-party inference providers (Cerebras, Together, Groq, Replicate, etc.) through a unified API
    
    **Key patterns covered:**
    
    - InferenceClient initialization and configuration
    - Chat Completion API (OpenAI-compatible messages format, streaming)
    - Text generation (raw completion, streaming)
    - Embeddings via feature extraction
    - Image generation (text-to-image)
    - Audio transcription (automatic speech recognition)
    - Translation, summarization, and text classification
    - Inference Endpoints (dedicated deployments)
    - Inference Providers (routing through third-party services)
    - Error handling with typed error classes
    
    **When NOT to use:**
    
    - If you only use OpenAI models -- use the OpenAI SDK directly
    - If you need a provider-agnostic unified SDK with structured outputs and tool calling -- use a higher-level AI SDK
    - If you need to fine-tune or train models -- use the `@huggingface/hub` package or Python `transformers`
    
    ---
    
    ## Examples Index
    
    - [Core: Setup, Chat & Text Generation](examples/core.md) -- Client init, chat completion, text generation, streaming, error handling
    - [Tasks: Embeddings, Vision, Audio & NLP](examples/tasks.md) -- Feature extraction, image generation, speech recognition, translation, summarization, classification
    - [Quick API Reference](reference.md) -- Method signatures, error types, provider list, model recommendations
    
    ---
    
    <philosophy>
    
    ## Philosophy
    
    The `@huggingface/inference` SDK provides a **unified TypeScript client** for accessing hundreds of thousands of ML models through multiple backends: serverless Inference Providers, dedicated Inference Endpoints, and local servers.
    
    **Core principles:**
    
    1. **Model-agnostic access** -- One client, any model on the Hub. Swap models by changing the `model` parameter without code changes.
    2. **Provider flexibility** -- Route inference through 20+ providers (Cerebras, Together, Groq, Replicate, etc.) with a single `provider` parameter, or deploy your own Inference Endpoints.
    3. **Task-oriented API** -- Methods map to ML tasks (`chatCompletion`, `textToImage`, `automaticSpeechRecognition`), not raw HTTP endpoints.
    4. **OpenAI-compatible chat** -- `chatCompletion()` uses the OpenAI message format (`role` + `content`), making migration between providers easy.
    5. **Streaming as async generators** -- `chatCompletionStream()` and `textGenerationStream()` return `AsyncGenerator`, consumed with `for await...of`.
    
    </philosophy>
    
    ---
    
    <patterns>
    
    ## Core Patterns
    
    ### Pattern 1: Client Setup
    
    Initialize with your Hugging Face access token. The token is required for authenticated access.
    
    ```typescript
    // lib/hf-client.ts -- basic setup
    import { InferenceClient } from "@huggingface/inference";
    
    const client = new InferenceClient(process.env.HF_TOKEN);
    
    export { client };
    ```
    
    ```typescript
    // lib/hf-client.ts -- with custom endpoint
    const ENDPOINT_URL =
      "https://your-endpoint.us-east-1.aws.endpoints.huggingface.cloud/v1/";
    
    const client = new InferenceClient(process.env.HF_TOKEN, {
      endpointUrl: ENDPOINT_URL,
    });
    
    export { client };
    ```
    
    **Why good:** Token from env var, named constant for endpoint URL, named export
    
    ```typescript
    // BAD: Hardcoded token, no named export
    const hf = new InferenceClient("hf_abc123xyz");
    export default hf;
    ```
    
    **Why bad:** Hardcoded token is a security risk, default export violates conventions
    
    **See:** [examples/core.md](examples/core.md) for provider routing, local endpoints, and endpoint helper
    
    ---
    
    ### Pattern 2: Chat Completion (OpenAI-Compatible)
    
    Use `chatCompletion()` for conversational LLM tasks. Follows the OpenAI message format.
    
    ```typescript
    const MAX_TOKENS = 512;
    const TEMPERATURE = 0.1;
    
    const response = await client.chatCompletion({
      model: "Qwen/Qwen3-32B",
      provider: "cerebras",
      messages: [
        { role: "system", content: "You are a helpful coding assistant." },
        { role: "user", content: "Explain TypeScript generics." },
      ],
      max_tokens: MAX_TOKENS,
      temperature: TEMPERATURE,
    });
    
    console.log(response.choices[0].message.content);
    ```
    
    **Why good:** Named constants for parameters, explicit model and provider, system message for behavior
    
    ```typescript
    // BAD: No model specified, magic numbers, no system message
    const response = await client.chatCompletion({
      messages: [{ role: "user", content: "do something" }],
      max_tokens: 512,
      temperature: 0.1,
    });
    ```
    
    **Why bad:** Missing required `model`, magic numbers, vague prompt, no system instruction
    
    **See:** [examples/core.md](examples/core.md) for multi-turn conversations and provider selection
    
    ---
    
    ### Pattern 3: Streaming Chat Completion
    
    Use `chatCompletionStream()` for streaming responses. Returns an `AsyncGenerator`.
    
    ```typescript
    const MAX_TOKENS = 512;
    let fullResponse = "";
    
    for await (const chunk of client.chatCompletionStream({
      model: "Qwen/Qwen3-32B",
      provider: "cerebras",
      messages: [{ role: "user", content: "Explain async/await in TypeScript." }],
      max_tokens: MAX_TOKENS,
    })) {
      if (chunk.choices && chunk.choices.length > 0) {
        const content = chunk.choices[0].delta.content;
        if (content) {
          process.stdout.write(content);
          fullResponse += content;
        }
      }
    }
    console.log(); // newline
    ```
    
    **Why good:** Async generator consumed with `for await`, progressive output, null checks on chunk data
    
    ```typescript
    // BAD: Not checking chunk.choices, ignoring null content
    for await (const chunk of client.chatCompletionStream({
      model: "...",
      messages: [],
    })) {
      process.stdout.write(chunk.choices[0].delta.content); // May throw on null
    }
    ```
    
    **Why bad:** No null check -- `choices` may be empty, `content` may be null between chunks
    
    **See:** [examples/core.md](examples/core.md) for text generation streaming
    
    ---
    
    ### Pattern 4: Text Generation (Raw Completion)
    
    Use `textGeneration()` for prompt continuation without the chat message format.
    
    ```typescript
    const MAX_NEW_TOKENS = 250;
    
    const result = await client.textGeneration({
      model: "mistralai/Mixtral-8x7B-v0.1",
      provider: "together",
      inputs: "The key benefits of TypeScript are",
      parameters: { max_new_tokens: MAX_NEW_TOKENS },
    });
    
    console.log(result.generated_text);
    ```
    
    **Why good:** Named constant, clear prompt, explicit provider, direct access to `generated_text`
    
    **See:** [examples/core.md](examples/core.md) for streaming text generation
    
    ---
    
    ### Pattern 5: Embeddings (Feature Extraction)
    
    Use `featureExtraction()` for generating vector embeddings for semantic search and RAG.
    
    ```typescript
    const embeddings = await client.featureExtraction({
      model: "sentence-transformers/all-MiniLM-L6-v2",
      inputs: "That is a happy person",
    });
    // Returns: number[] (embedding vector)
    ```
    
    **Why good:** Purpose-built embedding model, simple input/output
    
    **See:** [examples/tasks.md](examples/tasks.md) for batch embeddings and cosine similarity
    
    ---
    
    ### Pattern 6: Image Generation (Text-to-Image)
    
    Use `textToImage()` to generate images from text prompts. Returns a `Blob`.
    
    ```typescript
    const imageBlob = await client.textToImage({
      model: "black-forest-labs/FLUX.1-dev",
      inputs: "a serene mountain landscape at sunset",
      provider: "replicate",
    });
    // imageBlob is a Blob -- write to file or convert to buffer
    ```
    
    **Why good:** Explicit model and provider, descriptive prompt
    
    **See:** [examples/tasks.md](examples/tasks.md) for saving images, image-to-image, and output formats
    
    ---
    
    ### Pattern 7: Audio Transcription
    
    Use `automaticSpeechRecognition()` for speech-to-text.
    
    ```typescript
    import { readFileSync } from "node:fs";
    
    const result = await client.automaticSpeechRecognition({
      model: "facebook/wav2vec2-large-960h-lv60-self",
      data: readFileSync("audio/recording.flac"),
    });
    
    console.log(result.text);
    ```
    
    **Why good:** Uses `data` parameter with file buffer, outputs `.text`
    
    **See:** [examples/tasks.md](examples/tasks.md) for Whisper models, audio classification, and text-to-speech
    
    ---
    
    ### Pattern 8: Error Handling
    
    Always catch `InferenceClientError` and its subclasses. Re-throw unexpected errors.
    
    ```typescript
    import {
      InferenceClientError,
      InferenceClientInputError,
      InferenceClientProviderApiError,
      InferenceClientProviderOutputError,
      InferenceClientHubApiError,
    } from "@huggingface/inference";
    
    try {
      const result = await client.chatCompletion({
        model: "Qwen/Qwen3-32B",
        messages: [{ role: "user", content: "Hello" }],
      });
    } catch (error) {
      if (error instanceof InferenceClientProviderApiError) {
        console.error("Provider API error:", error.message);
        console.error("Request:", error.request);
        console.error("Response:", error.response);
      } else if (error instanceof InferenceClientHubApiError) {
        console.error("Hub API error:", error.message);
      } else if (error instanceof InferenceClientProviderOutputError) {
        console.error("Malformed provider response:", error.message);
      } else if (error instanceof InferenceClientInputError) {
        console.error("Invalid input:", error.message);
      } else if (error instanceof InferenceClientError) {
        console.error("Inference error:", error.message);
      } else {
        throw error; // Re-throw non-inference errors
      }
    }
    ```
    
    **Why good:** Specific error types for each failure mode, request/response details for debugging, re-throws unexpected errors
    
    **See:** [examples/core.md](examples/core.md) for full error handling patterns
    
    </patterns>
    
    ---
    
    <decision_framework>
    
    ## Decision Framework
    
    ### Which Method to Use
    
    ```
    What is your task?
    +-- Conversational LLM (messages) -> chatCompletion() / chatCompletionStream()
    +-- Raw text continuation -> textGeneration() / textGenerationStream()
    +-- Embeddings for search/RAG -> featureExtraction()
    +-- Image from text prompt -> textToImage()
    +-- Speech to text -> automaticSpeechRecognition()
    +-- Text to speech -> textToSpeech()
    +-- Language translation -> translation()
    +-- Summarize long text -> summarization()
    +-- Classify text -> textClassification()
    +-- Named entity recognition -> tokenClassification()
    +-- Classify image -> imageClassification()
    +-- Detect objects -> objectDetection()
    +-- Caption an image -> imageToText()
    +-- Answer questions from context -> questionAnswering()
    ```
    
    ### Chat Completion vs Text Generation
    
    ```
    Do you have a conversation with roles (system/user/assistant)?
    +-- YES -> chatCompletion() / chatCompletionStream()
    |   Uses OpenAI-compatible message format
    |   Supports system messages, multi-turn
    +-- NO -> Do you want to continue/complete a text prompt?
        +-- YES -> textGeneration() / textGenerationStream()
        |   Takes raw text input via 'inputs'
        +-- NO -> Use a task-specific method instead
    ```
    
    ### Serverless vs Dedicated
    
    ```
    What are your deployment needs?
    +-- Prototyping / low volume -> Serverless Inference Providers (provider: "auto")
    |   Free tier available, shared infrastructure, may have cold starts
    +-- Production / high volume -> Inference Endpoints (endpointUrl)
    |   Dedicated GPU, autoscaling, scale-to-zero, private infrastructure
    +-- Local development -> Local endpoint (endpointUrl: "http://localhost:8080")
        Works with llama.cpp, Ollama, vLLM, TGI, LiteLLM
    ```
    
    ### When to Use This SDK vs Others
    
    ```
    Do you need access to 200k+ open-source models?
    +-- YES -> Use @huggingface/inference
    +-- NO -> Do you only use OpenAI models?
        +-- YES -> Not this skill's scope -- use the OpenAI SDK directly
        +-- NO -> Do you need structured outputs / tool calling?
            +-- YES -> Not this skill's scope -- use a higher-level AI SDK
            +-- NO -> @huggingface/inference works for most ML tasks
    ```
    
    </decision_framework>
    
    ---
    
    <red_flags>
    
    ## RED FLAGS
    
    **High Priority Issues:**
    
    - Hardcoding access tokens instead of using environment variables (security breach risk)
    - Using bare `catch` blocks without checking `InferenceClientError` types (hides API errors, loses debug info)
    - Omitting the `model` parameter -- always specify the model explicitly for predictable behavior (the SDK can pick a recommended model if omitted, but this is unreliable for production)
    - Not consuming `chatCompletionStream()` / `textGenerationStream()` generators (tokens are silently lost)
    - Using `textGeneration()` for conversational tasks instead of `chatCompletion()` (wrong API shape)
    
    **Medium Priority Issues:**
    
    - Not specifying `provider` when a specific provider is needed (default `"auto"` picks based on your HF settings, which may not be optimal)
    - Not checking `chunk.choices` length before accessing `chunk.choices[0]` in streaming (may throw on empty chunks)
    - Using `request()` / `streamingRequest()` directly -- these are deprecated, use task-specific methods
    - Ignoring `max_tokens` / `max_new_tokens` limits (output may be truncated or excessively long)
    - Not handling model loading time for serverless inference (cold models return 503, then load)
    
    **Common Mistakes:**
    
    - Confusing `chatCompletion()` parameters with `textGeneration()` parameters -- chat uses `messages` + `max_tokens`, text generation uses `inputs` + `parameters.max_new_tokens`
    - Using `inputs` parameter with `chatCompletion()` -- it uses `messages`, not `inputs`
    - Using `messages` parameter with `textGeneration()` -- it uses `inputs`, not `messages`
    - Forgetting that `textToImage()` returns a `Blob`, not a URL or Buffer
    - Treating `featureExtraction()` output as always a flat array -- shape depends on the model (can be nested arrays for batch inputs)
    - Not passing `data` (binary) for audio/image tasks, or passing a string path instead of the actual file buffer
    
    **Gotchas & Edge Cases:**
    
    - The `provider: "auto"` default selects providers based on your HF account settings at `hf.co/settings/inference-providers` -- not by availability or speed. Set an explicit provider for predictable routing.
    - Serverless models may need time to load (cold start). First requests to a cold model may return 503 errors while the model warms up. The SDK handles retries, but initial requests can be slow.
    - When using Inference Endpoints with `endpointUrl`, the model parameter is often ignored because the endpoint serves a specific model.
    - `chatCompletion()` is OpenAI-API compatible -- it works with any OpenAI-compatible endpoint, not just Hugging Face.
    - `HfInference` is still exported for backward compatibility but `InferenceClient` is the current class name.
    - Third-party provider API keys can be passed as the `accessToken` -- when authenticated with a non-HF key, requests go directly to the provider instead of through HF's routing layer.
    - Tree-shakeable imports (`import { textGeneration } from "@huggingface/inference"`) require passing `accessToken` as a parameter instead of constructor.
    - `textToImage()` supports multiple output types: `blob` (default), `url`, `dataUrl`, or `json` via the options `outputType` parameter.
    - Translation requires `parameters.src_lang` and `parameters.tgt_lang` for many-to-many models like `mbart-large-50-many-to-many-mmt`.
    
    </red_flags>
    
    ---
    
    <critical_reminders>
    
    ## CRITICAL REMINDERS
    
    > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
    
    **(You MUST always pass an access token to `InferenceClient` -- never deploy without authentication)**
    
    **(You MUST use `chatCompletion()` / `chatCompletionStream()` for conversational LLM tasks -- these follow the OpenAI-compatible message format)**
    
    **(You MUST handle errors using `InferenceClientError` and its subclasses -- never use bare catch blocks without error type checking)**
    
    **(You MUST specify a `model` parameter for every inference call -- there is no default model)**
    
    **(You MUST never hardcode access tokens -- always use environment variables via `process.env.HF_TOKEN`)**
    
    **Failure to follow these rules will produce insecure, unreliable, or silently failing AI integrations.**
    
    </critical_reminders>
    

Comments (0)

Sign in to join the conversation.

No comments yet.

Reviews (0)

No reviews yet.

Related