ai-infrastructure-huggingface-inference
Hugging Face Inference SDK patterns for TypeScript/Node.js — InferenceClient setup, chat completion, text generation, streaming, embeddings, image generation, audio transcription, translation, summarization, and Inference Endpoints
Install
npx skills add https://github.com/agents-inc/skills/tree/main/dist/plugins/ai-infrastructure-huggingface-inference/skills/ai-infrastructure-huggingface-inference
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install agents-inc-skills@llmmart
git clone https://github.com/agents-inc/skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole agents-inc/skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Hugging Face Inference Patterns
Quick Guide: Use
@huggingface/inference(v4+) to access 200k+ ML models on the Hugging Face Hub. UseInferenceClientwithchatCompletion()for OpenAI-compatible chat,textGeneration()for raw text completion,chatCompletionStream()for streaming,featureExtraction()for embeddings,textToImage()for image generation, andautomaticSpeechRecognition()for audio transcription. Setproviderto route through inference providers (Cerebras, Together, Groq, etc.) or useendpointUrlfor dedicated Inference Endpoints.
<critical_requirements>
CRITICAL: Before Using This Skill
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST always pass an access token to InferenceClient -- never deploy without authentication)
(You MUST use chatCompletion() / chatCompletionStream() for conversational LLM tasks -- these follow the OpenAI-compatible message format)
(You MUST handle errors using InferenceClientError and its subclasses -- never use bare catch blocks without error type checking)
(You MUST specify a model parameter for every inference call -- there is no default model)
(You MUST never hardcode access tokens -- always use environment variables via process.env.HF_TOKEN)
</critical_requirements>
Auto-detection: Hugging Face, huggingface, @huggingface/inference, InferenceClient, HfInference, hf.chatCompletion, hf.textGeneration, hf.featureExtraction, hf.textToImage, hf.automaticSpeechRecognition, hf.translation, hf.summarization, hf.textToSpeech, chatCompletionStream, textGenerationStream, HF_TOKEN, inference provider, Inference Endpoints
When to use:
- Accessing any of the 200k+ models hosted on the Hugging Face Hub
- Running chat completion with open-source LLMs (Qwen, Mistral, Llama, etc.)
- Generating embeddings with sentence-transformer models for semantic search
- Generating images from text prompts (FLUX, Stable Diffusion)
- Transcribing audio with automatic speech recognition models
- Running translation, summarization, text classification, or NER tasks
- Deploying models on dedicated Inference Endpoints for production use
- Using third-party inference providers (Cerebras, Together, Groq, Replicate, etc.) through a unified API
Key patterns covered:
- InferenceClient initialization and configuration
- Chat Completion API (OpenAI-compatible messages format, streaming)
- Text generation (raw completion, streaming)
- Embeddings via feature extraction
- Image generation (text-to-image)
- Audio transcription (automatic speech recognition)
- Translation, summarization, and text classification
- Inference Endpoints (dedicated deployments)
- Inference Providers (routing through third-party services)
- Error handling with typed error classes
When NOT to use:
- If you only use OpenAI models -- use the OpenAI SDK directly
- If you need a provider-agnostic unified SDK with structured outputs and tool calling -- use a higher-level AI SDK
- If you need to fine-tune or train models -- use the
@huggingface/hubpackage or Pythontransformers
Examples Index
- Core: Setup, Chat & Text Generation -- Client init, chat completion, text generation, streaming, error handling
- Tasks: Embeddings, Vision, Audio & NLP -- Feature extraction, image generation, speech recognition, translation, summarization, classification
- Quick API Reference -- Method signatures, error types, provider list, model recommendations
<decision_framework>
Decision Framework
Which Method to Use
What is your task?
+-- Conversational LLM (messages) -> chatCompletion() / chatCompletionStream()
+-- Raw text continuation -> textGeneration() / textGenerationStream()
+-- Embeddings for search/RAG -> featureExtraction()
+-- Image from text prompt -> textToImage()
+-- Speech to text -> automaticSpeechRecognition()
+-- Text to speech -> textToSpeech()
+-- Language translation -> translation()
+-- Summarize long text -> summarization()
+-- Classify text -> textClassification()
+-- Named entity recognition -> tokenClassification()
+-- Classify image -> imageClassification()
+-- Detect objects -> objectDetection()
+-- Caption an image -> imageToText()
+-- Answer questions from context -> questionAnswering()
Chat Completion vs Text Generation
Do you have a conversation with roles (system/user/assistant)?
+-- YES -> chatCompletion() / chatCompletionStream()
| Uses OpenAI-compatible message format
| Supports system messages, multi-turn
+-- NO -> Do you want to continue/complete a text prompt?
+-- YES -> textGeneration() / textGenerationStream()
| Takes raw text input via 'inputs'
+-- NO -> Use a task-specific method instead
Serverless vs Dedicated
What are your deployment needs?
+-- Prototyping / low volume -> Serverless Inference Providers (provider: "auto")
| Free tier available, shared infrastructure, may have cold starts
+-- Production / high volume -> Inference Endpoints (endpointUrl)
| Dedicated GPU, autoscaling, scale-to-zero, private infrastructure
+-- Local development -> Local endpoint (endpointUrl: "http://localhost:8080")
Works with llama.cpp, Ollama, vLLM, TGI, LiteLLM
When to Use This SDK vs Others
Do you need access to 200k+ open-source models?
+-- YES -> Use @huggingface/inference
+-- NO -> Do you only use OpenAI models?
+-- YES -> Not this skill's scope -- use the OpenAI SDK directly
+-- NO -> Do you need structured outputs / tool calling?
+-- YES -> Not this skill's scope -- use a higher-level AI SDK
+-- NO -> @huggingface/inference works for most ML tasks
</decision_framework>
<red_flags>
RED FLAGS
High Priority Issues:
- Hardcoding access tokens instead of using environment variables (security breach risk)
- Using bare
catchblocks without checkingInferenceClientErrortypes (hides API errors, loses debug info) - Omitting the
modelparameter -- always specify the model explicitly for predictable behavior (the SDK can pick a recommended model if omitted, but this is unreliable for production) - Not consuming
chatCompletionStream()/textGenerationStream()generators (tokens are silently lost) - Using
textGeneration()for conversational tasks instead ofchatCompletion()(wrong API shape)
Medium Priority Issues:
- Not specifying
providerwhen a specific provider is needed (default"auto"picks based on your HF settings, which may not be optimal) - Not checking
chunk.choiceslength before accessingchunk.choices[0]in streaming (may throw on empty chunks) - Using
request()/streamingRequest()directly -- these are deprecated, use task-specific methods - Ignoring
max_tokens/max_new_tokenslimits (output may be truncated or excessively long) - Not handling model loading time for serverless inference (cold models return 503, then load)
Common Mistakes:
- Confusing
chatCompletion()parameters withtextGeneration()parameters -- chat usesmessages+max_tokens, text generation usesinputs+parameters.max_new_tokens - Using
inputsparameter withchatCompletion()-- it usesmessages, notinputs - Using
messagesparameter withtextGeneration()-- it usesinputs, notmessages - Forgetting that
textToImage()returns aBlob, not a URL or Buffer - Treating
featureExtraction()output as always a flat array -- shape depends on the model (can be nested arrays for batch inputs) - Not passing
data(binary) for audio/image tasks, or passing a string path instead of the actual file buffer
Gotchas & Edge Cases:
- The
provider: "auto"default selects providers based on your HF account settings athf.co/settings/inference-providers-- not by availability or speed. Set an explicit provider for predictable routing. - Serverless models may need time to load (cold start). First requests to a cold model may return 503 errors while the model warms up. The SDK handles retries, but initial requests can be slow.
- When using Inference Endpoints with
endpointUrl, the model parameter is often ignored because the endpoint serves a specific model. chatCompletion()is OpenAI-API compatible -- it works with any OpenAI-compatible endpoint, not just Hugging Face.HfInferenceis still exported for backward compatibility butInferenceClientis the current class name.- Third-party provider API keys can be passed as the
accessToken-- when authenticated with a non-HF key, requests go directly to the provider instead of through HF's routing layer. - Tree-shakeable imports (
import { textGeneration } from "@huggingface/inference") require passingaccessTokenas a parameter instead of constructor. textToImage()supports multiple output types:blob(default),url,dataUrl, orjsonvia the optionsoutputTypeparameter.- Translation requires
parameters.src_langandparameters.tgt_langfor many-to-many models likembart-large-50-many-to-many-mmt.
</red_flags>
<critical_reminders>
CRITICAL REMINDERS
All code must follow project conventions in CLAUDE.md (kebab-case, named exports, import ordering,
import type, named constants)
(You MUST always pass an access token to InferenceClient -- never deploy without authentication)
(You MUST use chatCompletion() / chatCompletionStream() for conversational LLM tasks -- these follow the OpenAI-compatible message format)
(You MUST handle errors using InferenceClientError and its subclasses -- never use bare catch blocks without error type checking)
(You MUST specify a model parameter for every inference call -- there is no default model)
(You MUST never hardcode access tokens -- always use environment variables via process.env.HF_TOKEN)
Failure to follow these rules will produce insecure, unreliable, or silently failing AI integrations.
</critical_reminders>
Files (skills)
-
examples
-
core.md 10.3 KB
# Hugging Face Inference -- Setup, Chat & Text Generation Examples > Client initialization, chat completion, text generation, streaming, providers, endpoints, and error handling. See [SKILL.md](../SKILL.md) for core patterns. **Related examples:** - [tasks.md](tasks.md) -- Embeddings, image generation, audio, translation, summarization --- ## Basic Client Setup ```typescript // lib/hf-client.ts import { InferenceClient } from "@huggingface/inference"; // Requires HF_TOKEN environment variable const client = new InferenceClient(process.env.HF_TOKEN); export { client }; ``` --- ## With Specific Provider ```typescript // lib/hf-client.ts import { InferenceClient } from "@huggingface/inference"; // HF token routes requests through HF's proxy to the provider const client = new InferenceClient(process.env.HF_TOKEN); // Provider specified per-request const response = await client.chatCompletion({ model: "Qwen/Qwen3-32B", provider: "cerebras", // Route to Cerebras infrastructure messages: [{ role: "user", content: "Hello!" }], }); ``` --- ## With Dedicated Inference Endpoint ```typescript // lib/hf-endpoint.ts import { InferenceClient } from "@huggingface/inference"; // Option 1: endpointUrl in constructor const ENDPOINT_URL = "https://j3z5luu0ooo76jnl.us-east-1.aws.endpoints.huggingface.cloud/v1/"; const client = new InferenceClient(process.env.HF_TOKEN, { endpointUrl: ENDPOINT_URL, }); // Model parameter is ignored -- endpoint serves a specific model const response = await client.chatCompletion({ messages: [{ role: "user", content: "What is the capital of France?" }], }); console.log(response.choices[0].message.content); export { client }; ``` ```typescript // Option 2: .endpoint() helper import { InferenceClient } from "@huggingface/inference"; const ENDPOINT_URL = "https://j3z5luu0ooo76jnl.us-east-1.aws.endpoints.huggingface.cloud/v1/"; const client = new InferenceClient(process.env.HF_TOKEN); const endpointClient = client.endpoint(ENDPOINT_URL); const response = await endpointClient.chatCompletion({ messages: [{ role: "user", content: "Hello!" }], }); ``` --- ## With Local Endpoint (Ollama, llama.cpp, vLLM, TGI, LiteLLM) ```typescript // lib/hf-local.ts import { InferenceClient } from "@huggingface/inference"; // No token needed for local endpoints const client = new InferenceClient(undefined, { endpointUrl: "http://localhost:8080", }); const response = await client.chatCompletion({ messages: [{ role: "user", content: "What is the capital of France?" }], }); console.log(response.choices[0].message.content); ``` --- ## Disable Endpoint Retry-on-Error ```typescript // By default, calls wait for model to load (scale-to-zero endpoints) // Disable to handle 503 errors yourself const response = await client.chatCompletion( { messages: [{ role: "user", content: "Hello" }], }, { retry_on_error: false, }, ); ``` --- ## Chat Completion -- Basic ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const MAX_TOKENS = 512; const TEMPERATURE = 0.1; async function chat(userMessage: string): Promise<string> { const response = await client.chatCompletion({ model: "Qwen/Qwen3-32B", provider: "cerebras", messages: [ { role: "system", content: "You are a helpful assistant. Be concise." }, { role: "user", content: userMessage }, ], max_tokens: MAX_TOKENS, temperature: TEMPERATURE, }); const content = response.choices[0].message.content; if (!content) { throw new Error("No content in response"); } return content; } const answer = await chat("What is TypeScript in one sentence?"); console.log(answer); ``` --- ## Chat Completion -- Multi-Turn ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const MAX_TOKENS = 512; interface Message { role: "system" | "user" | "assistant"; content: string; } const messages: Message[] = [ { role: "system", content: "You are a TypeScript expert." }, { role: "user", content: "What is a union type?" }, ]; const response = await client.chatCompletion({ model: "Qwen/Qwen3-32B", provider: "cerebras", messages, max_tokens: MAX_TOKENS, }); // Append assistant response for next turn const assistantContent = response.choices[0].message.content ?? ""; messages.push({ role: "assistant", content: assistantContent }); messages.push({ role: "user", content: "Give me a real-world example." }); const followUp = await client.chatCompletion({ model: "Qwen/Qwen3-32B", provider: "cerebras", messages, max_tokens: MAX_TOKENS, }); ``` --- ## Chat Completion -- Streaming ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const MAX_TOKENS = 512; async function streamChat(prompt: string): Promise<string> { let fullResponse = ""; for await (const chunk of client.chatCompletionStream({ model: "Qwen/Qwen3-32B", provider: "cerebras", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: prompt }, ], max_tokens: MAX_TOKENS, })) { if (chunk.choices && chunk.choices.length > 0) { const content = chunk.choices[0].delta.content; if (content) { process.stdout.write(content); fullResponse += content; } } } console.log(); // newline return fullResponse; } const result = await streamChat("Explain promises in JavaScript."); console.log("Response length:", result.length); ``` --- ## Text Generation -- Non-Streaming ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const MAX_NEW_TOKENS = 250; const result = await client.textGeneration({ model: "mistralai/Mixtral-8x7B-v0.1", provider: "together", inputs: "The key benefits of TypeScript are", parameters: { max_new_tokens: MAX_NEW_TOKENS }, }); console.log(result.generated_text); ``` --- ## Text Generation -- Streaming ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const MAX_NEW_TOKENS = 250; for await (const output of client.textGenerationStream({ model: "mistralai/Mixtral-8x7B-v0.1", provider: "together", inputs: "The key benefits of TypeScript are", parameters: { max_new_tokens: MAX_NEW_TOKENS }, })) { if (output.token.text) { process.stdout.write(output.token.text); } } console.log(); ``` --- ## Tree-Shakeable Imports ```typescript // Import individual functions for smaller bundles import { textGeneration, chatCompletion } from "@huggingface/inference"; const MAX_NEW_TOKENS = 250; // Each function requires accessToken as a parameter const result = await textGeneration({ accessToken: process.env.HF_TOKEN, model: "mistralai/Mixtral-8x7B-v0.1", provider: "together", inputs: "The key benefits of TypeScript are", parameters: { max_new_tokens: MAX_NEW_TOKENS }, }); ``` --- ## Third-Party Provider API Key (Direct) ```typescript import { InferenceClient } from "@huggingface/inference"; // Using a provider's own API key bypasses HF routing -- goes directly to provider const client = new InferenceClient(process.env.MISTRAL_API_KEY, { endpointUrl: "https://api.mistral.ai", }); const response = await client.chatCompletion({ model: "mistral-tiny", messages: [{ role: "user", content: "Hello!" }], }); ``` --- ## Error Handling -- Production Pattern ```typescript import { InferenceClient } from "@huggingface/inference"; import { InferenceClientError, InferenceClientInputError, InferenceClientProviderApiError, InferenceClientProviderOutputError, InferenceClientHubApiError, } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const MAX_TOKENS = 512; async function safeChat(prompt: string): Promise<string | null> { try { const response = await client.chatCompletion({ model: "Qwen/Qwen3-32B", provider: "cerebras", messages: [ { role: "system", content: "You are a helpful assistant." }, { role: "user", content: prompt }, ], max_tokens: MAX_TOKENS, }); return response.choices[0].message.content; } catch (error) { if (error instanceof InferenceClientProviderApiError) { // API-level errors from the inference provider (rate limits, auth, server errors) console.error("Provider API error:", error.message); console.error("Request details:", error.request); console.error("Response details:", error.response); return null; } if (error instanceof InferenceClientHubApiError) { // Errors from the HF Hub API (model not found, repo issues) console.error("Hub API error:", error.message); return null; } if (error instanceof InferenceClientProviderOutputError) { // Provider returned malformed response console.error("Malformed provider response:", error.message); return null; } if (error instanceof InferenceClientInputError) { // Invalid input parameters console.error("Invalid input:", error.message); return null; } // Catch-all for any @huggingface/inference error if (error instanceof InferenceClientError) { console.error("Inference error:", error.message); return null; } // Unknown errors should be re-thrown throw error; } } const result = await safeChat("Hello!"); if (result) { console.log(result); } else { console.error("Failed to get response"); } ``` --- ## Error Type Reference ```typescript // Error class hierarchy: // InferenceClientError (base) // +-- InferenceClientInputError (invalid input parameters) // +-- InferenceClientProviderApiError (provider API errors: rate limits, auth, 5xx) // | .request -- Request details (URL, method, headers) // | .response -- Response details (status code, body) // +-- InferenceClientHubApiError (HF Hub API errors: model not found) // | .request -- Request details // | .response -- Response details // +-- InferenceClientProviderOutputError (malformed provider response) ``` --- _For task-specific examples (embeddings, images, audio), see [tasks.md](tasks.md). For API reference tables, see [reference.md](../reference.md)._ -
tasks.md 10.9 KB
# Hugging Face Inference -- Task Examples (Embeddings, Vision, Audio, NLP) > Task-specific examples: feature extraction, image generation, speech recognition, translation, summarization, classification, and more. See [SKILL.md](../SKILL.md) for core patterns. **Prerequisites**: Understand client setup and chat completion from [core.md](core.md) first. **Related examples:** - [core.md](core.md) -- Client setup, chat completion, text generation, streaming --- ## Feature Extraction (Embeddings) ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); // Single input -- returns number[] const embedding = await client.featureExtraction({ model: "sentence-transformers/all-MiniLM-L6-v2", inputs: "That is a happy person", }); ``` --- ## Batch Embeddings with Cosine Similarity ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); function cosineSimilarity(a: number[], b: number[]): number { let dotProduct = 0; let normA = 0; let normB = 0; for (let i = 0; i < a.length; i++) { dotProduct += a[i] * b[i]; normA += a[i] * a[i]; normB += b[i] * b[i]; } return dotProduct / (Math.sqrt(normA) * Math.sqrt(normB)); } const texts = [ "TypeScript adds static types to JavaScript", "JavaScript is a dynamic programming language", "The weather is nice today", ]; // featureExtraction accepts a single string input // For batch, call individually or use a model that supports array inputs const embeddings = await Promise.all( texts.map((text) => client.featureExtraction({ model: "sentence-transformers/all-MiniLM-L6-v2", inputs: text, }), ), ); // Compare first text to all others const SIMILARITY_THRESHOLD = 0.5; for (let i = 1; i < texts.length; i++) { const similarity = cosineSimilarity( embeddings[0] as number[], embeddings[i] as number[], ); const isRelated = similarity > SIMILARITY_THRESHOLD; console.log( `"${texts[0]}" vs "${texts[i]}": ${similarity.toFixed(4)} (${isRelated ? "related" : "unrelated"})`, ); } ``` --- ## Text-to-Image ```typescript import { InferenceClient } from "@huggingface/inference"; import { writeFileSync } from "node:fs"; const client = new InferenceClient(process.env.HF_TOKEN); // Default output is Blob const imageBlob = await client.textToImage({ model: "black-forest-labs/FLUX.1-dev", inputs: "a serene mountain landscape at sunset, oil painting style", provider: "replicate", }); // Write Blob to file const buffer = Buffer.from(await imageBlob.arrayBuffer()); writeFileSync("output/landscape.png", buffer); ``` --- ## Text-to-Image with Output Options ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); // Get as URL string const imageUrl = await client.textToImage( { model: "black-forest-labs/FLUX.1-dev", inputs: "a cat wearing a hat", }, { outputType: "url" }, ); console.log("Image URL:", imageUrl); // Get as data URL (base64 embedded) const dataUrl = await client.textToImage( { model: "black-forest-labs/FLUX.1-dev", inputs: "a cat wearing a hat", }, { outputType: "dataUrl" }, ); // Use dataUrl directly in <img src="..."> ``` --- ## Image Classification ```typescript import { InferenceClient } from "@huggingface/inference"; import { readFileSync } from "node:fs"; const client = new InferenceClient(process.env.HF_TOKEN); const results = await client.imageClassification({ model: "google/vit-base-patch16-224", data: readFileSync("images/photo.png"), }); // Results: Array<{ label: string, score: number }> for (const result of results) { console.log(`${result.label}: ${(result.score * 100).toFixed(1)}%`); } ``` --- ## Image to Text (Captioning) ```typescript import { InferenceClient } from "@huggingface/inference"; import { readFileSync } from "node:fs"; const client = new InferenceClient(process.env.HF_TOKEN); const result = await client.imageToText({ model: "nlpconnect/vit-gpt2-image-captioning", data: readFileSync("images/photo.png"), }); console.log("Caption:", result.generated_text); ``` --- ## Object Detection ```typescript import { InferenceClient } from "@huggingface/inference"; import { readFileSync } from "node:fs"; const client = new InferenceClient(process.env.HF_TOKEN); const detections = await client.objectDetection({ model: "facebook/detr-resnet-50", data: readFileSync("images/street.png"), }); // Results: Array<{ label: string, score: number, box: { xmin, ymin, xmax, ymax } }> for (const detection of detections) { console.log( `${detection.label} (${(detection.score * 100).toFixed(1)}%) at [${detection.box.xmin}, ${detection.box.ymin}]`, ); } ``` --- ## Automatic Speech Recognition (Transcription) ```typescript import { InferenceClient } from "@huggingface/inference"; import { readFileSync } from "node:fs"; const client = new InferenceClient(process.env.HF_TOKEN); const result = await client.automaticSpeechRecognition({ model: "facebook/wav2vec2-large-960h-lv60-self", data: readFileSync("audio/recording.flac"), }); console.log("Transcript:", result.text); ``` --- ## Audio Classification ```typescript import { InferenceClient } from "@huggingface/inference"; import { readFileSync } from "node:fs"; const client = new InferenceClient(process.env.HF_TOKEN); const results = await client.audioClassification({ model: "superb/hubert-large-superb-er", data: readFileSync("audio/sample.flac"), }); // Results: Array<{ label: string, score: number }> for (const result of results) { console.log(`${result.label}: ${(result.score * 100).toFixed(1)}%`); } ``` --- ## Text to Speech ```typescript import { InferenceClient } from "@huggingface/inference"; import { writeFileSync } from "node:fs"; const client = new InferenceClient(process.env.HF_TOKEN); const audioBlob = await client.textToSpeech({ model: "espnet/kan-bayashi_ljspeech_vits", inputs: "Hello, welcome to the application!", }); const buffer = Buffer.from(await audioBlob.arrayBuffer()); writeFileSync("output/speech.wav", buffer); ``` --- ## Translation ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); // Simple translation (model determines source/target) const result = await client.translation({ model: "t5-base", inputs: "My name is Wolfgang and I live in Berlin", }); console.log("Translation:", result.translation_text); ``` ```typescript // Many-to-many model with explicit language codes const result = await client.translation({ model: "facebook/mbart-large-50-many-to-many-mmt", inputs: "My name is Wolfgang and I live in Berlin", parameters: { src_lang: "en_XX", tgt_lang: "fr_XX", }, }); console.log("French:", result.translation_text); ``` --- ## Summarization ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const MAX_LENGTH = 100; const result = await client.summarization({ model: "facebook/bart-large-cnn", inputs: "The tower is 324 metres (1,063 ft) tall, about the same height as an 81-storey building, " + "and the tallest structure in Paris. Its base is square, measuring 125 metres (410 ft) on " + "each side. During its construction, the Eiffel Tower surpassed the Washington Monument to " + "become the tallest man-made structure in the world, a title it held for 41 years until the " + "Chrysler Building in New York City was finished in 1930.", parameters: { max_length: MAX_LENGTH }, }); console.log("Summary:", result.summary_text); ``` --- ## Text Classification (Sentiment Analysis) ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const results = await client.textClassification({ model: "distilbert-base-uncased-finetuned-sst-2-english", inputs: "I love this product! It works perfectly.", }); // Results: Array<{ label: string, score: number }> for (const result of results) { console.log(`${result.label}: ${(result.score * 100).toFixed(1)}%`); } ``` --- ## Token Classification (Named Entity Recognition) ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const entities = await client.tokenClassification({ model: "dbmdz/bert-large-cased-finetuned-conll03-english", inputs: "My name is Sarah Jessica Parker but you can call me Jessica", }); // Results: Array<{ entity_group: string, word: string, score: number, start: number, end: number }> for (const entity of entities) { console.log( `${entity.word} -> ${entity.entity_group} (${(entity.score * 100).toFixed(1)}%)`, ); } ``` --- ## Zero-Shot Classification ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const result = await client.zeroShotClassification({ model: "facebook/bart-large-mnli", inputs: [ "Hi, I recently bought a device from your company but it is not working as advertised and I would like to get reimbursed!", ], parameters: { candidate_labels: ["refund", "legal", "faq"] }, }); // Result includes labels sorted by score console.log(result); ``` --- ## Question Answering (Extractive) ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const result = await client.questionAnswering({ model: "deepset/roberta-base-squad2", inputs: { question: "What is the capital of France?", context: "The capital of France is Paris. It is located in the north of the country.", }, }); console.log(`Answer: ${result.answer} (score: ${result.score.toFixed(4)})`); ``` --- ## Sentence Similarity ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const result = await client.sentenceSimilarity({ model: "sentence-transformers/paraphrase-xlm-r-multilingual-v1", inputs: { source_sentence: "That is a happy person", sentences: [ "That is a happy dog", "That is a very happy person", "Today is a sunny day", ], }, }); // Returns: number[] (similarity scores for each sentence) console.log("Similarity scores:", result); ``` --- ## Visual Question Answering ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); const result = await client.visualQuestionAnswering({ model: "dandelin/vilt-b32-finetuned-vqa", inputs: { question: "How many cats are in the image?", image: await (await fetch("https://example.com/cats.jpg")).blob(), }, }); console.log(`Answer: ${result.answer} (score: ${result.score.toFixed(4)})`); ``` --- _For client setup and chat patterns, see [core.md](core.md). For API reference tables, see [reference.md](../reference.md)._
-
-
reference.md 10.9 KB
# Hugging Face Inference Quick Reference > Method signatures, error types, provider list, and model recommendations. See [SKILL.md](SKILL.md) for core concepts and [examples/](examples/) for code examples. --- ## Package Installation ```bash npm install @huggingface/inference ``` --- ## Client Initialization ```typescript import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(accessToken, { endpointUrl?: string, // Custom endpoint URL (Inference Endpoints or local) }); ``` ### Environment Variables | Variable | Purpose | | ---------- | ------------------------------------ | | `HF_TOKEN` | Hugging Face access token (required) | --- ## API Methods Reference ### NLP -- Chat & Text Generation ```typescript // Chat Completion (OpenAI-compatible) const response = await client.chatCompletion({ model: string, // Required: model ID on the Hub messages: Message[], // Required: { role, content }[] max_tokens?: number, // Max output tokens temperature?: number, // 0-2 (default: 1) top_p?: number, // Nucleus sampling provider?: string, // Inference provider (default: "auto") }); // Returns: { choices: [{ message: { role, content } }] } // Chat Completion Streaming const stream = client.chatCompletionStream({ model: string, messages: Message[], max_tokens?: number, temperature?: number, provider?: string, }); // Returns: AsyncGenerator<{ choices: [{ delta: { content } }] }> // Text Generation const result = await client.textGeneration({ model: string, // Required inputs: string, // Required: prompt text parameters?: { max_new_tokens?: number, // Max tokens to generate temperature?: number, top_p?: number, repetition_penalty?: number, }, provider?: string, }); // Returns: { generated_text: string } // Text Generation Streaming const stream = client.textGenerationStream({ model: string, inputs: string, parameters?: { ... }, provider?: string, }); // Returns: AsyncGenerator<{ token: { text, id, logprob }, generated_text?: string }> ``` ### NLP -- Analysis Tasks ```typescript // Feature Extraction (Embeddings) await client.featureExtraction({ model: string, // Required inputs: string, // Required: text to embed }); // Returns: number[] | number[][] // Summarization await client.summarization({ model: string, inputs: string, parameters?: { max_length?: number, min_length?: number }, }); // Returns: { summary_text: string } // Translation await client.translation({ model: string, inputs: string, parameters?: { src_lang?: string, tgt_lang?: string }, }); // Returns: { translation_text: string } // Text Classification await client.textClassification({ model: string, inputs: string, }); // Returns: Array<{ label: string, score: number }> // Token Classification (NER) await client.tokenClassification({ model: string, inputs: string, }); // Returns: Array<{ entity_group: string, word: string, score: number, start: number, end: number }> // Zero-Shot Classification await client.zeroShotClassification({ model: string, inputs: string[], parameters: { candidate_labels: string[] }, }); // Question Answering await client.questionAnswering({ model: string, inputs: { question: string, context: string }, }); // Returns: { answer: string, score: number, start: number, end: number } // Fill Mask await client.fillMask({ model: string, inputs: string, // Must contain [MASK] token }); // Sentence Similarity await client.sentenceSimilarity({ model: string, inputs: { source_sentence: string, sentences: string[] }, }); // Returns: number[] ``` ### Audio ```typescript // Automatic Speech Recognition await client.automaticSpeechRecognition({ model: string, data: Blob | ArrayBuffer, // Audio file data }); // Returns: { text: string } // Audio Classification await client.audioClassification({ model: string, data: Blob | ArrayBuffer, }); // Returns: Array<{ label: string, score: number }> // Text to Speech await client.textToSpeech({ model: string, inputs: string, }); // Returns: Blob (audio) // Audio to Audio await client.audioToAudio({ model: string, data: Blob | ArrayBuffer, }); // Returns: Array<{ blob: Blob, label: string, "content-type": string }> ``` ### Computer Vision ```typescript // Text to Image await client.textToImage({ model: string, inputs: string, // Text prompt provider?: string, }, { outputType?: "blob" | "url" | "dataUrl" | "json", // Default: "blob" }); // Returns: Blob | string (depends on outputType) // Image Classification await client.imageClassification({ model: string, data: Blob | ArrayBuffer, }); // Returns: Array<{ label: string, score: number }> // Object Detection await client.objectDetection({ model: string, data: Blob | ArrayBuffer, }); // Returns: Array<{ label: string, score: number, box: { xmin, ymin, xmax, ymax } }> // Image to Text (Captioning) await client.imageToText({ model: string, data: Blob | ArrayBuffer, }); // Returns: { generated_text: string } // Image Segmentation await client.imageSegmentation({ model: string, data: Blob | ArrayBuffer, }); // Returns: Array<{ label: string, score: number, mask: string }> // Image to Image await client.imageToImage({ inputs: Blob, parameters?: { prompt?: string }, model: string, }); // Returns: Blob // Zero-Shot Image Classification await client.zeroShotImageClassification({ model: string, inputs: { image: Blob }, parameters: { candidate_labels: string[] }, }); ``` ### Multimodal ```typescript // Visual Question Answering await client.visualQuestionAnswering({ model: string, inputs: { question: string, image: Blob }, }); // Returns: { answer: string, score: number } // Document Question Answering await client.documentQuestionAnswering({ model: string, inputs: { question: string, image: Blob }, }); ``` --- ## Recommended Models by Task | Task | Model | Notes | | -------------------------- | -------------------------------------------------- | ------------------------ | | Chat Completion | `Qwen/Qwen3-32B` | Strong open-source LLM | | Chat Completion | `mistralai/Mixtral-8x7B-v0.1` | Mixture of experts | | Text Generation | `mistralai/Mixtral-8x7B-v0.1` | Text continuation | | Embeddings | `sentence-transformers/all-MiniLM-L6-v2` | Fast, good quality | | Multilingual Embeddings | `intfloat/multilingual-e5-large` | Cross-lingual | | Text-to-Image | `black-forest-labs/FLUX.1-dev` | High quality image gen | | Image Classification | `google/vit-base-patch16-224` | General purpose | | Object Detection | `facebook/detr-resnet-50` | General purpose | | Speech Recognition | `facebook/wav2vec2-large-960h-lv60-self` | English | | Audio Classification | `superb/hubert-large-superb-er` | Emotion recognition | | Text-to-Speech | `espnet/kan-bayashi_ljspeech_vits` | English TTS | | Translation | `t5-base` | English focused | | Translation (multilingual) | `facebook/mbart-large-50-many-to-many-mmt` | 50 languages | | Summarization | `facebook/bart-large-cnn` | News summarization | | Sentiment Analysis | `distilbert-base-uncased-finetuned-sst-2-english` | Binary sentiment | | NER | `dbmdz/bert-large-cased-finetuned-conll03-english` | Named entities | | Zero-Shot Classification | `facebook/bart-large-mnli` | No training needed | | Question Answering | `deepset/roberta-base-squad2` | Extractive QA | | Image Captioning | `nlpconnect/vit-gpt2-image-captioning` | General captioning | | Fill Mask | `bert-base-uncased` | Masked language modeling | | Zero-Shot Image Class. | `openai/clip-vit-large-patch14-336` | No training needed | --- ## Error Types | Error Class | Cause | Has `.request` / `.response`? | | ------------------------------------ | -------------------------------------------- | ----------------------------- | | `InferenceClientError` | Base class for all errors | No | | `InferenceClientInputError` | Invalid input parameters | No | | `InferenceClientProviderApiError` | Provider API errors (rate limits, auth, 5xx) | Yes | | `InferenceClientHubApiError` | HF Hub API errors (model not found) | Yes | | `InferenceClientProviderOutputError` | Malformed provider response | No | All errors extend the base `Error` class via `InferenceClientError`. --- ## Supported Inference Providers | Provider | Key Tasks Supported | | --------------- | -------------------------------- | | Baseten | Chat, text generation | | Blackforestlabs | Image generation | | Cerebras | Chat completion, text generation | | Clarifai | Chat, text, image generation | | Cohere | Chat, text, embeddings | | DeepInfra | Chat, text, embeddings | | Fal.ai | Image, video generation | | Featherless AI | Chat, text generation | | Fireworks AI | Chat, text generation | | Groq | Chat completion (fast inference) | | HF Inference | All tasks (Hugging Face's own) | | Hyperbolic | Chat, text generation | | Nebius | Chat, text generation | | Novita | Chat, image generation | | Nscale | Chat, text generation | | NVIDIA | Chat, text, embeddings | | OVHcloud | Chat, text generation | | Public AI | Chat, text generation | | Replicate | Image generation, audio | | Sambanova | Chat, text generation | | Scaleway | Chat, text generation | | Together | Chat, text, image generation | | Wavespeed.ai | Image, video generation | | Z.ai | Chat, text generation | Set `provider: "auto"` (default) to use your account's preferred provider order. --- See [examples/core.md](examples/core.md) for tree-shakeable imports and endpoint configuration patterns. -
SKILL.md 18 KB
--- name: ai-infrastructure-huggingface-inference description: Hugging Face Inference SDK patterns for TypeScript/Node.js — InferenceClient setup, chat completion, text generation, streaming, embeddings, image generation, audio transcription, translation, summarization, and Inference Endpoints --- # Hugging Face Inference Patterns > **Quick Guide:** Use `@huggingface/inference` (v4+) to access 200k+ ML models on the Hugging Face Hub. Use `InferenceClient` with `chatCompletion()` for OpenAI-compatible chat, `textGeneration()` for raw text completion, `chatCompletionStream()` for streaming, `featureExtraction()` for embeddings, `textToImage()` for image generation, and `automaticSpeechRecognition()` for audio transcription. Set `provider` to route through inference providers (Cerebras, Together, Groq, etc.) or use `endpointUrl` for dedicated Inference Endpoints. --- <critical_requirements> ## CRITICAL: Before Using This Skill > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST always pass an access token to `InferenceClient` -- never deploy without authentication)** **(You MUST use `chatCompletion()` / `chatCompletionStream()` for conversational LLM tasks -- these follow the OpenAI-compatible message format)** **(You MUST handle errors using `InferenceClientError` and its subclasses -- never use bare catch blocks without error type checking)** **(You MUST specify a `model` parameter for every inference call -- there is no default model)** **(You MUST never hardcode access tokens -- always use environment variables via `process.env.HF_TOKEN`)** </critical_requirements> --- **Auto-detection:** Hugging Face, huggingface, @huggingface/inference, InferenceClient, HfInference, hf.chatCompletion, hf.textGeneration, hf.featureExtraction, hf.textToImage, hf.automaticSpeechRecognition, hf.translation, hf.summarization, hf.textToSpeech, chatCompletionStream, textGenerationStream, HF_TOKEN, inference provider, Inference Endpoints **When to use:** - Accessing any of the 200k+ models hosted on the Hugging Face Hub - Running chat completion with open-source LLMs (Qwen, Mistral, Llama, etc.) - Generating embeddings with sentence-transformer models for semantic search - Generating images from text prompts (FLUX, Stable Diffusion) - Transcribing audio with automatic speech recognition models - Running translation, summarization, text classification, or NER tasks - Deploying models on dedicated Inference Endpoints for production use - Using third-party inference providers (Cerebras, Together, Groq, Replicate, etc.) through a unified API **Key patterns covered:** - InferenceClient initialization and configuration - Chat Completion API (OpenAI-compatible messages format, streaming) - Text generation (raw completion, streaming) - Embeddings via feature extraction - Image generation (text-to-image) - Audio transcription (automatic speech recognition) - Translation, summarization, and text classification - Inference Endpoints (dedicated deployments) - Inference Providers (routing through third-party services) - Error handling with typed error classes **When NOT to use:** - If you only use OpenAI models -- use the OpenAI SDK directly - If you need a provider-agnostic unified SDK with structured outputs and tool calling -- use a higher-level AI SDK - If you need to fine-tune or train models -- use the `@huggingface/hub` package or Python `transformers` --- ## Examples Index - [Core: Setup, Chat & Text Generation](examples/core.md) -- Client init, chat completion, text generation, streaming, error handling - [Tasks: Embeddings, Vision, Audio & NLP](examples/tasks.md) -- Feature extraction, image generation, speech recognition, translation, summarization, classification - [Quick API Reference](reference.md) -- Method signatures, error types, provider list, model recommendations --- <philosophy> ## Philosophy The `@huggingface/inference` SDK provides a **unified TypeScript client** for accessing hundreds of thousands of ML models through multiple backends: serverless Inference Providers, dedicated Inference Endpoints, and local servers. **Core principles:** 1. **Model-agnostic access** -- One client, any model on the Hub. Swap models by changing the `model` parameter without code changes. 2. **Provider flexibility** -- Route inference through 20+ providers (Cerebras, Together, Groq, Replicate, etc.) with a single `provider` parameter, or deploy your own Inference Endpoints. 3. **Task-oriented API** -- Methods map to ML tasks (`chatCompletion`, `textToImage`, `automaticSpeechRecognition`), not raw HTTP endpoints. 4. **OpenAI-compatible chat** -- `chatCompletion()` uses the OpenAI message format (`role` + `content`), making migration between providers easy. 5. **Streaming as async generators** -- `chatCompletionStream()` and `textGenerationStream()` return `AsyncGenerator`, consumed with `for await...of`. </philosophy> --- <patterns> ## Core Patterns ### Pattern 1: Client Setup Initialize with your Hugging Face access token. The token is required for authenticated access. ```typescript // lib/hf-client.ts -- basic setup import { InferenceClient } from "@huggingface/inference"; const client = new InferenceClient(process.env.HF_TOKEN); export { client }; ``` ```typescript // lib/hf-client.ts -- with custom endpoint const ENDPOINT_URL = "https://your-endpoint.us-east-1.aws.endpoints.huggingface.cloud/v1/"; const client = new InferenceClient(process.env.HF_TOKEN, { endpointUrl: ENDPOINT_URL, }); export { client }; ``` **Why good:** Token from env var, named constant for endpoint URL, named export ```typescript // BAD: Hardcoded token, no named export const hf = new InferenceClient("hf_abc123xyz"); export default hf; ``` **Why bad:** Hardcoded token is a security risk, default export violates conventions **See:** [examples/core.md](examples/core.md) for provider routing, local endpoints, and endpoint helper --- ### Pattern 2: Chat Completion (OpenAI-Compatible) Use `chatCompletion()` for conversational LLM tasks. Follows the OpenAI message format. ```typescript const MAX_TOKENS = 512; const TEMPERATURE = 0.1; const response = await client.chatCompletion({ model: "Qwen/Qwen3-32B", provider: "cerebras", messages: [ { role: "system", content: "You are a helpful coding assistant." }, { role: "user", content: "Explain TypeScript generics." }, ], max_tokens: MAX_TOKENS, temperature: TEMPERATURE, }); console.log(response.choices[0].message.content); ``` **Why good:** Named constants for parameters, explicit model and provider, system message for behavior ```typescript // BAD: No model specified, magic numbers, no system message const response = await client.chatCompletion({ messages: [{ role: "user", content: "do something" }], max_tokens: 512, temperature: 0.1, }); ``` **Why bad:** Missing required `model`, magic numbers, vague prompt, no system instruction **See:** [examples/core.md](examples/core.md) for multi-turn conversations and provider selection --- ### Pattern 3: Streaming Chat Completion Use `chatCompletionStream()` for streaming responses. Returns an `AsyncGenerator`. ```typescript const MAX_TOKENS = 512; let fullResponse = ""; for await (const chunk of client.chatCompletionStream({ model: "Qwen/Qwen3-32B", provider: "cerebras", messages: [{ role: "user", content: "Explain async/await in TypeScript." }], max_tokens: MAX_TOKENS, })) { if (chunk.choices && chunk.choices.length > 0) { const content = chunk.choices[0].delta.content; if (content) { process.stdout.write(content); fullResponse += content; } } } console.log(); // newline ``` **Why good:** Async generator consumed with `for await`, progressive output, null checks on chunk data ```typescript // BAD: Not checking chunk.choices, ignoring null content for await (const chunk of client.chatCompletionStream({ model: "...", messages: [], })) { process.stdout.write(chunk.choices[0].delta.content); // May throw on null } ``` **Why bad:** No null check -- `choices` may be empty, `content` may be null between chunks **See:** [examples/core.md](examples/core.md) for text generation streaming --- ### Pattern 4: Text Generation (Raw Completion) Use `textGeneration()` for prompt continuation without the chat message format. ```typescript const MAX_NEW_TOKENS = 250; const result = await client.textGeneration({ model: "mistralai/Mixtral-8x7B-v0.1", provider: "together", inputs: "The key benefits of TypeScript are", parameters: { max_new_tokens: MAX_NEW_TOKENS }, }); console.log(result.generated_text); ``` **Why good:** Named constant, clear prompt, explicit provider, direct access to `generated_text` **See:** [examples/core.md](examples/core.md) for streaming text generation --- ### Pattern 5: Embeddings (Feature Extraction) Use `featureExtraction()` for generating vector embeddings for semantic search and RAG. ```typescript const embeddings = await client.featureExtraction({ model: "sentence-transformers/all-MiniLM-L6-v2", inputs: "That is a happy person", }); // Returns: number[] (embedding vector) ``` **Why good:** Purpose-built embedding model, simple input/output **See:** [examples/tasks.md](examples/tasks.md) for batch embeddings and cosine similarity --- ### Pattern 6: Image Generation (Text-to-Image) Use `textToImage()` to generate images from text prompts. Returns a `Blob`. ```typescript const imageBlob = await client.textToImage({ model: "black-forest-labs/FLUX.1-dev", inputs: "a serene mountain landscape at sunset", provider: "replicate", }); // imageBlob is a Blob -- write to file or convert to buffer ``` **Why good:** Explicit model and provider, descriptive prompt **See:** [examples/tasks.md](examples/tasks.md) for saving images, image-to-image, and output formats --- ### Pattern 7: Audio Transcription Use `automaticSpeechRecognition()` for speech-to-text. ```typescript import { readFileSync } from "node:fs"; const result = await client.automaticSpeechRecognition({ model: "facebook/wav2vec2-large-960h-lv60-self", data: readFileSync("audio/recording.flac"), }); console.log(result.text); ``` **Why good:** Uses `data` parameter with file buffer, outputs `.text` **See:** [examples/tasks.md](examples/tasks.md) for Whisper models, audio classification, and text-to-speech --- ### Pattern 8: Error Handling Always catch `InferenceClientError` and its subclasses. Re-throw unexpected errors. ```typescript import { InferenceClientError, InferenceClientInputError, InferenceClientProviderApiError, InferenceClientProviderOutputError, InferenceClientHubApiError, } from "@huggingface/inference"; try { const result = await client.chatCompletion({ model: "Qwen/Qwen3-32B", messages: [{ role: "user", content: "Hello" }], }); } catch (error) { if (error instanceof InferenceClientProviderApiError) { console.error("Provider API error:", error.message); console.error("Request:", error.request); console.error("Response:", error.response); } else if (error instanceof InferenceClientHubApiError) { console.error("Hub API error:", error.message); } else if (error instanceof InferenceClientProviderOutputError) { console.error("Malformed provider response:", error.message); } else if (error instanceof InferenceClientInputError) { console.error("Invalid input:", error.message); } else if (error instanceof InferenceClientError) { console.error("Inference error:", error.message); } else { throw error; // Re-throw non-inference errors } } ``` **Why good:** Specific error types for each failure mode, request/response details for debugging, re-throws unexpected errors **See:** [examples/core.md](examples/core.md) for full error handling patterns </patterns> --- <decision_framework> ## Decision Framework ### Which Method to Use ``` What is your task? +-- Conversational LLM (messages) -> chatCompletion() / chatCompletionStream() +-- Raw text continuation -> textGeneration() / textGenerationStream() +-- Embeddings for search/RAG -> featureExtraction() +-- Image from text prompt -> textToImage() +-- Speech to text -> automaticSpeechRecognition() +-- Text to speech -> textToSpeech() +-- Language translation -> translation() +-- Summarize long text -> summarization() +-- Classify text -> textClassification() +-- Named entity recognition -> tokenClassification() +-- Classify image -> imageClassification() +-- Detect objects -> objectDetection() +-- Caption an image -> imageToText() +-- Answer questions from context -> questionAnswering() ``` ### Chat Completion vs Text Generation ``` Do you have a conversation with roles (system/user/assistant)? +-- YES -> chatCompletion() / chatCompletionStream() | Uses OpenAI-compatible message format | Supports system messages, multi-turn +-- NO -> Do you want to continue/complete a text prompt? +-- YES -> textGeneration() / textGenerationStream() | Takes raw text input via 'inputs' +-- NO -> Use a task-specific method instead ``` ### Serverless vs Dedicated ``` What are your deployment needs? +-- Prototyping / low volume -> Serverless Inference Providers (provider: "auto") | Free tier available, shared infrastructure, may have cold starts +-- Production / high volume -> Inference Endpoints (endpointUrl) | Dedicated GPU, autoscaling, scale-to-zero, private infrastructure +-- Local development -> Local endpoint (endpointUrl: "http://localhost:8080") Works with llama.cpp, Ollama, vLLM, TGI, LiteLLM ``` ### When to Use This SDK vs Others ``` Do you need access to 200k+ open-source models? +-- YES -> Use @huggingface/inference +-- NO -> Do you only use OpenAI models? +-- YES -> Not this skill's scope -- use the OpenAI SDK directly +-- NO -> Do you need structured outputs / tool calling? +-- YES -> Not this skill's scope -- use a higher-level AI SDK +-- NO -> @huggingface/inference works for most ML tasks ``` </decision_framework> --- <red_flags> ## RED FLAGS **High Priority Issues:** - Hardcoding access tokens instead of using environment variables (security breach risk) - Using bare `catch` blocks without checking `InferenceClientError` types (hides API errors, loses debug info) - Omitting the `model` parameter -- always specify the model explicitly for predictable behavior (the SDK can pick a recommended model if omitted, but this is unreliable for production) - Not consuming `chatCompletionStream()` / `textGenerationStream()` generators (tokens are silently lost) - Using `textGeneration()` for conversational tasks instead of `chatCompletion()` (wrong API shape) **Medium Priority Issues:** - Not specifying `provider` when a specific provider is needed (default `"auto"` picks based on your HF settings, which may not be optimal) - Not checking `chunk.choices` length before accessing `chunk.choices[0]` in streaming (may throw on empty chunks) - Using `request()` / `streamingRequest()` directly -- these are deprecated, use task-specific methods - Ignoring `max_tokens` / `max_new_tokens` limits (output may be truncated or excessively long) - Not handling model loading time for serverless inference (cold models return 503, then load) **Common Mistakes:** - Confusing `chatCompletion()` parameters with `textGeneration()` parameters -- chat uses `messages` + `max_tokens`, text generation uses `inputs` + `parameters.max_new_tokens` - Using `inputs` parameter with `chatCompletion()` -- it uses `messages`, not `inputs` - Using `messages` parameter with `textGeneration()` -- it uses `inputs`, not `messages` - Forgetting that `textToImage()` returns a `Blob`, not a URL or Buffer - Treating `featureExtraction()` output as always a flat array -- shape depends on the model (can be nested arrays for batch inputs) - Not passing `data` (binary) for audio/image tasks, or passing a string path instead of the actual file buffer **Gotchas & Edge Cases:** - The `provider: "auto"` default selects providers based on your HF account settings at `hf.co/settings/inference-providers` -- not by availability or speed. Set an explicit provider for predictable routing. - Serverless models may need time to load (cold start). First requests to a cold model may return 503 errors while the model warms up. The SDK handles retries, but initial requests can be slow. - When using Inference Endpoints with `endpointUrl`, the model parameter is often ignored because the endpoint serves a specific model. - `chatCompletion()` is OpenAI-API compatible -- it works with any OpenAI-compatible endpoint, not just Hugging Face. - `HfInference` is still exported for backward compatibility but `InferenceClient` is the current class name. - Third-party provider API keys can be passed as the `accessToken` -- when authenticated with a non-HF key, requests go directly to the provider instead of through HF's routing layer. - Tree-shakeable imports (`import { textGeneration } from "@huggingface/inference"`) require passing `accessToken` as a parameter instead of constructor. - `textToImage()` supports multiple output types: `blob` (default), `url`, `dataUrl`, or `json` via the options `outputType` parameter. - Translation requires `parameters.src_lang` and `parameters.tgt_lang` for many-to-many models like `mbart-large-50-many-to-many-mmt`. </red_flags> --- <critical_reminders> ## CRITICAL REMINDERS > **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants) **(You MUST always pass an access token to `InferenceClient` -- never deploy without authentication)** **(You MUST use `chatCompletion()` / `chatCompletionStream()` for conversational LLM tasks -- these follow the OpenAI-compatible message format)** **(You MUST handle errors using `InferenceClientError` and its subclasses -- never use bare catch blocks without error type checking)** **(You MUST specify a `model` parameter for every inference call -- there is no default model)** **(You MUST never hardcode access tokens -- always use environment variables via `process.env.HF_TOKEN`)** **Failure to follow these rules will produce insecure, unreliable, or silently failing AI integrations.** </critical_reminders>
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.