Claude
Skill
dotnet-microsoft-extensions-ai
Build provider-agnostic .NET AI integrations with `Microsoft.Extensions.AI`, `IChatClient`, embeddings, middleware, structured output, vector search, and evaluation.
Virus-scanned
Reviewed automatically before listing.
Download
postpartum-genushyacinthus29-dotnet-skills-skills_dotnet-microsoft-extensions-ai-bfa4ebd.zip · 133 KB
Install
skills CLI
npx skills add https://github.com/Postpartum-genushyacinthus29/dotnet-skills/tree/main/skills/dotnet-microsoft-extensions-ai
Claude Code
claude plugin marketplace add https://llmmart.ai/marketplace.json && claude plugin install postpartum-genushyacinthus29-dotnet-skills@llmmart
Git
git clone https://github.com/Postpartum-genushyacinthus29/dotnet-skills.git
The skills CLI installs just this skill, for any of its supported agents. Claude Code installs the whole postpartum-genushyacinthus29/dotnet-skills collection as a plugin from our marketplace. Git is the plain clone.
Skill manifest
Microsoft.Extensions.AI
Trigger On
- building or reviewing
.NETcode that usesMicrosoft.Extensions.AI,Microsoft.Extensions.AI.Abstractions,IChatClient,IEmbeddingGenerator,ChatOptions, orAIFunction - adding
IImageGenerator, local-model chat via Ollama, AI app templates, or the.NET AIquickstarts for assistants and MCP - choosing between low-level AI abstractions, provider SDKs, vector-search composition, evaluation libraries, and a fuller agent framework
- adding streaming chat, structured output, embeddings, tool calling, telemetry, caching, or DI-based AI middleware
- wiring
Microsoft.Extensions.VectorData,Microsoft.Extensions.DataIngestion, MCP tooling, or evaluation packages around a provider-agnostic AI app
Workflow
- Classify the request first: plain model access, tool calling, embeddings/vector search, evaluation, image generation, local-model prototyping, MCP bootstrap, or true agent orchestration.
- Default to
Microsoft.Extensions.AIfor application and service code that needs provider-agnostic chat, embeddings, middleware, structured output, and testability. - Reference
Microsoft.Extensions.AI.Abstractionsdirectly only when authoring provider libraries or lower-level reusable integration packages. - Model
IChatClientandIEmbeddingGeneratorcomposition explicitly in DI. Keep options, caching, telemetry, logging, and tool invocation inspectable in the pipeline. - Treat chat state deliberately. For stateless providers, resend history. For stateful providers, propagate
ConversationIdrather than assuming all providers behave the same way. - Use
Microsoft.Extensions.VectorDataandMicrosoft.Extensions.DataIngestionas adjacent building blocks for RAG instead of hand-rolling store abstractions prematurely. Model ingestion as an explicit reader -> processor -> chunker -> writer pipeline when the document-preparation path matters. - Treat the
.NET AIquickstarts as bootstrap paths, not finished architecture. They now cover minimal assistants, MCP client/server flows, local models, app templates, and image generation. Start there for a vertical slice, then harden the DI, telemetry, and evaluation story here. - Escalate to
dotnet-microsoft-agent-frameworkwhen the requirement becomes agent threads, multi-agent orchestration, higher-order workflows, durable execution, or remote agent hosting. - Validate with real providers, realistic prompts, and evaluation gates so the abstraction layer actually buys portability and reliability.
Architecture
flowchart LR
A["Task"] --> B{"Need agent threads, multi-agent orchestration, or remote agent hosting?"}
B -->|Yes| C["Use Microsoft Agent Framework on top of `Microsoft.Extensions.AI.Abstractions`"]
B -->|No| D{"Need provider-agnostic chat, embeddings, tools, typed output, or evaluation?"}
D -->|Yes| E["Use `Microsoft.Extensions.AI`"]
E --> F["Compose `IChatClient` / `IEmbeddingGenerator` in DI"]
F --> G["Add caching, telemetry, tools, vector data, and evaluation deliberately"]
D -->|No| H["Use plain provider SDKs or deterministic .NET code"]
Core Knowledge
Microsoft.Extensions.AI.Abstractionscontains the core exchange contracts such asIChatClient,IEmbeddingGenerator<TInput, TEmbedding>, message/content types, and tool abstractions.Microsoft.Extensions.AIadds the higher-level application surface: middleware builders, automatic function invocation, caching, logging, and OpenTelemetry integration.- Most apps and services should reference
Microsoft.Extensions.AI; provider and connector libraries usually reference only the abstractions package. IChatClientcenters onGetResponseAsyncandGetStreamingResponseAsync. The returnedChatResponseorChatResponseUpdateobjects carry messages, tool-related content, metadata, and optional conversation identifiers.- Local-model quickstarts still route through the same
IChatClientabstraction. Ollama-backed clients are useful for low-cost prototyping, offline dev loops, and portability testing, but you still own chat history replay, latency, and model-quality tradeoffs. ChatOptionsis the normal control plane for model ID, temperature, tools,AdditionalProperties, and provider-specific raw options.- Tool calling is modeled with
AIFunction,AIFunctionFactory, andFunctionInvokingChatClient. Ambient data can flow through closures,AdditionalProperties,AIFunctionArguments.Context, or DI. - Tool calling can target local .NET methods, external APIs, or MCP-backed tools. The model requests calls; your app still owns execution, validation, and side-effect boundaries.
- Tool definitions consume request tokens. Keep tool descriptions short and register only the tools relevant for the current conversation or workflow.
FunctionInvokingChatClientcan handle the tool-invocation loop and parallel tool-call responses automatically when the provider/model supports that shape.IEmbeddingGeneratoris the standard abstraction for semantic search, vector indexing, similarity, and cache-key generation. Pair it withMicrosoft.Extensions.VectorData.Abstractionsfor vector store operations.IImageGeneratoris the experimental MEAI image surface. TreatMEAI001as an intentional opt-in, keep image generation separate from chat concerns, and compose logging/caching/hosting middleware around it the same way you would forIChatClient.Microsoft.Extensions.DataIngestiongives you the document-side RAG pipeline:IngestionDocument, document readers like MarkItDown/Markdig, document processors such asImageAlternativeTextEnricher, chunkers, chunk processors,VectorStoreWriter<T>, andIngestionPipeline<T>for end-to-end composition.IngestionPipeline<T>.ProcessAsyncis partial-success oriented. HandleIAsyncEnumerable<IngestionResult>deliberately instead of assuming one failed document should automatically crash the whole ingestion run.Microsoft.Extensions.AI.Evaluation.*gives you quality, NLP, safety, caching, and reporting layers for regression checks and CI gates.- The official
.NET AIdocs now make MCP, assistants, local models, templates, and text-to-image part of the same app-level story. Usedotnet-mcpwhen the protocol itself becomes the design problem; stay here when you still mostly need app composition aroundIChatClientand friends. Microsoft Agent Frameworkbuilds on these abstractions. Use it when you need autonomous orchestration, threads, workflows, hosting, or multi-agent collaboration instead of just model composition.
Decision Cheatsheet
| If you need | Default choice | Why |
|---|---|---|
| App-level provider abstraction with middleware | Microsoft.Extensions.AI |
Highest leverage for apps and services |
| A reusable provider or connector library | Microsoft.Extensions.AI.Abstractions |
Keeps your package at the contract layer |
| Typed chat or UI streaming | IChatClient with GetResponseAsync / GetStreamingResponseAsync |
Common request/response shape across providers |
| Tool calling from .NET methods | AIFunction + FunctionInvokingChatClient |
Native function metadata and invocation pipeline |
| Typed structured output | IChatClient.GetResponseAsync<T> extensions |
Keeps schema intent in code instead of prompt parsing |
| Vector search or RAG | IEmbeddingGenerator + Microsoft.Extensions.VectorData.Abstractions |
Standardizes embeddings and store access |
| Local model prototyping | IChatClient with an Ollama-backed implementation |
Keeps the app on the MEAI abstractions while you validate prompts or UX locally |
| Text-to-image or image-generation middleware | IImageGenerator |
Use the dedicated image abstraction instead of overloading chat APIs |
| Evaluation and regression gates | Microsoft.Extensions.AI.Evaluation.* |
Relevance, safety, task adherence, caching, reports |
| Agent threads or multi-step autonomous orchestration | dotnet-microsoft-agent-framework |
This is beyond plain provider abstraction |
Common Failure Modes
- Referencing only
Microsoft.Extensions.AI.Abstractionsin an app and then rebuilding middleware, telemetry, or function invocation by hand. - Treating
IChatClientas if it already gives you durable agent threads, orchestration, or hosted-agent semantics. - Mixing provider-specific assistants APIs with
IChatClientas if they were the same runtime contract. - Forgetting to distinguish stateless history replay from stateful
ConversationIdflows. - Hiding important chat behavior in singleton service fields instead of explicit message history, options, or persistent storage.
- Adding tool calling without validating parameter binding, invalid input behavior, side effects, or DI-scoped dependencies.
- Building RAG without stable chunking, embedding-model/version tracking, or vector dimension discipline.
- Shipping AI features without evaluation baselines, safety checks, or telemetry for prompt/model drift.
Deliver
- a justified package and abstraction choice:
Abstractionsonly vs fullMicrosoft.Extensions.AI - a concrete
IChatClient/IEmbeddingGeneratorcomposition strategy - explicit tool-calling, options, state, caching, logging, and telemetry decisions
- vector-search, evaluation, or MCP integration guidance when the scenario needs it
- a clear escalation path to Agent Framework when the problem exceeds provider abstraction
Validate
- the abstraction layer solves a real portability, testability, or composition problem
- provider registration and middleware order stay explicit in DI
- chat state management matches whether the provider is stateless or stateful
- structured output, tool invocation, and embedding flows are typed and observable
- vector store, embedding model, and chunking strategy are consistent
- evaluation or safety gates exist for important prompts and agent-like behaviors
- agentic requirements are not being under-modeled as a simple
IChatClientintegration
When exact wording, edge-case API behavior, or less-common examples matter, check the local official docs snapshot before relying on summaries.
References
- official-docs-index.md - Slim local snapshot map with direct links to every mirrored
.NET AIdocs page plus API-reference pointers - patterns.md - Package choice,
IChatClient, embeddings, DI pipelines, tool-calling, and Agent Framework escalation guidance - examples.md - Quickstart-to-task map covering chat, structured output, function calling, vector search, local models, MCP, and assistants
- evaluation.md - Quality, NLP, safety, caching, reporting, and CI-oriented evaluation guidance
Files (dotnet-skills)
-
references
-
official-docs
-
conceptual
-
agents.md 3.1 KB
--- title: Agents description: Introduction to agents author: luisquintanilla ms.author: luquinta ms.date: 12/10/2025 ms.topic: concept-article --- # Agents This article introduces the core concepts behind agents, why they matter, and how they fit into workflows, setting you up to get started building agents in .NET. ## What are agents? **Agents are systems that accomplish objectives.** Agents become more capable when equipped with the following: - **Reasoning and decision-making**: Powered by LLMs, search algorithms, or planning and decision-making systems. - **Tool usage**: Access to Model Context Protocol (MCP) servers, code execution, and external APIs. - **Context awareness**: Informed by chat history, threads, vector stores, enterprise data, or knowledge graphs. These capabilities allow agents to operate more autonomously, adaptively, and intelligently. ## What are workflows? As objectives grow in complexity, they need to be broken down into manageable steps. That's where workflows come in. **Workflows define the sequence of steps required to achieve an objective.** Imagine you're launching a new feature on your business website. If it's a simple update, you might go from idea to production in a few hours. But for more complex initiatives, the process might include: - Requirement gathering - Design and architecture - Implementation - Testing - Deployment A few important observations: - Each step might contain subtasks. - Different specialists might own different phases. - Progress isn’t always linear. Bugs found during testing might send you back to implementation. - Success depends on planning, orchestration, and communication across stakeholders. ### Agents + workflows = agentic workflows Workflows don't require agents, but agents can supercharge them. When agents are equipped with reasoning, tools, and context, they can optimize workflows. This is the foundation of multi-agent systems, where agents collaborate within workflows to achieve complex goals. ### Workflow orchestration Agentic workflows can be orchestrated in a variety of ways. The following are a few of the most common: - [Sequential](#sequential) - [Concurrent](#concurrent) - [Handoff](#handoff) - [Group chat](#group-chat) - [Magentic](#magentic) #### Sequential Agents process tasks one after another, passing results forward. #### Concurrent Agents work in parallel, each handling different aspects of the task. #### Handoff Responsibility shifts from one agent to another based on conditions or outcomes. #### Group chat Agents collaborate in a shared conversation, exchanging insights in real-time. #### Magentic A lead agent directs other agents. ## How can I get started building agents in .NET? The building blocks in <xref:Microsoft.Extensions.AI> and <xref:Microsoft.Extensions.VectorData> supply the foundations for agents by providing modular components for AI models, tools, and data. These components serve as the foundation for Microsoft Agent Framework. For more information, see [Microsoft Agent Framework](/agent-framework/overview/agent-framework-overview). -
ai-tools.md 5.5 KB
--- title: "AI tool calling" description: "Understand how tool calling lets you integrate external tools with AI models across providers using Microsoft.Extensions.AI." ms.topic: concept-article ms.date: 03/03/2026 ai-usage: ai-assisted --- # AI tool calling *Tool calling* is an AI model capability that lets you describe available tools to an AI model so the model can request that your application invoke them. Tools can be .NET methods, calls to external APIs, interactions with [Model Context Protocol (MCP)](../get-started-mcp.md) servers, or any other executable operation. Instead of directly executing those tools, the model returns a structured output describing which tools to call and with what arguments. Your application invokes those tools and sends the results back to the model, enabling it to build a more accurate and grounded response. <xref:Microsoft.Extensions.AI> (MEAI) provides provider-agnostic abstractions for tool calling that work across AI services, including Azure OpenAI, OpenAI, Ollama, and others. You write your tool-calling logic once, and it works regardless of which underlying model or provider you use. ## Why use tool calling Tool calling simplifies how you connect external tools to AI models. You describe each tool to the model as part of the conversation. The model then decides which tools to invoke based on the user's question. After your application invokes the requested tools and returns the results, the model uses those results to construct a more complete and accurate response. Common use cases for tool calling include: - Answering questions by calling external APIs. For example, checking the weather forecast, or sending email. - Retrieving information from internal data stores. For example, aggregating sales data to answer, "What are my best-selling products?" - Producing structured data from unstructured text. For example, constructing a user profile from chat history. ## Call AI functions in MEAI The general flow for calling AI functions with <xref:Microsoft.Extensions.AI.IChatClient> is: 1. Define .NET methods as functions and configure them on a <xref:Microsoft.Extensions.AI.ChatOptions> instance. 1. Send the user's message to the model. The model decides which functions, if any, to call. It returns a structured response that lists the function calls and their arguments. > [!NOTE] > Models might hallucinate arguments that weren't described in your function definitions. 1. Parse the model's response and invoke the requested functions with the specified arguments. 1. Send another request that includes the function results as new messages in the conversation history. 1. The model responds with more function call requests or a final answer to the user's question. Continue invoking requested functions until the model provides a final response. MEAI's <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient> handles steps 3 through 5 automatically, so you don't need to manage the invocation loop yourself. ## Key types MEAI provides the following types to support function calling: - <xref:Microsoft.Extensions.AI.AIFunction>: Represents a function that can be described to an AI model, and invoked. This is the core abstraction for a function in MEAI. - <xref:Microsoft.Extensions.AI.AIFunctionFactory>: Provides factory methods for creating `AIFunction` instances from .NET methods. Use `AIFunctionFactory` to wrap existing methods as functions without writing boilerplate description or argument-parsing code. - <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient>: Wraps any `IChatClient` and adds automatic function-invocation capabilities. When the model requests a function call, `FunctionInvokingChatClient` invokes the corresponding `AIFunction`, collects the result, and continues the conversation—all transparently. ## Parallel function calling Some models support *parallel function calling*, where the model requests multiple function invocations in a single response. Your application invokes each function and returns all results together in one follow-up message. Parallel function calling reduces the number of round trips to the model, which lowers latency and API usage. `FunctionInvokingChatClient` supports parallel function calling automatically. ## Cross-provider support One of the key benefits of using MEAI for function calling is provider independence. The `AIFunction`, `AIFunctionFactory`, and `FunctionInvokingChatClient` types work with any `IChatClient` implementation, including: - Azure OpenAI - OpenAI - Ollama - Any other provider that implements `IChatClient` Because function calling support varies across models and providers, check your provider's documentation to confirm whether a specific model supports function calling or parallel function calling. ## Token considerations Tool descriptions are included in the request sent to the model and count against the model's token limit. This means tool definitions contribute to both token consumption and request cost. If your request approaches the model's token limit, consider these adjustments: - Reduce the number of tools registered for the conversation. - Shorten the method names and descriptions used to generate tool definitions. - Limit tool registration to only the tools relevant for a given conversation context. ## Related content - [Invoke .NET functions using an AI model](../quickstarts/use-function-calling.md) - [Use the IChatClient interface](../ichatclient.md) - [Understanding tokens](understanding-tokens.md) - [Prompt engineering](prompt-engineering-dotnet.md) -
chain-of-thought-prompting.md 2.7 KB
--- title: "Chain of thought prompting - .NET" description: "Learn how chain of thought prompting can simplify prompt engineering." ms.topic: concept-article #Don't change. ms.date: 03/04/2026 ai-usage: ai-assisted #customer intent: As a .NET developer, I want to understand what chain-of-thought prompting is and how it can help me save time and get better completions out of prompt engineering. --- # Chain of thought prompting GPT model performance and response quality benefit from *prompt engineering*, which is the practice of providing instructions and examples to a model to prime or refine its output. As they process instructions, models make more reasoning errors when they try to answer right away rather than taking time to work out an answer. Help the model reason its way toward correct answers more reliably by asking the model to include its chain of thought—that is, the steps it took to follow an instruction, along with the results of each step. *Chain of thought prompting* is the practice of prompting a model to perform a task step-by-step and to present each step and its result in order in the output. This simplifies prompt engineering by offloading some execution planning to the model, and makes it easier to connect any problem to a specific step so you know where to focus further efforts. It's generally simpler to instruct the model to include its chain of thought, but you can also use examples to show the model how to break down tasks. The following sections show both ways. ## Use chain of thought prompting in instructions To use an instruction for chain of thought prompting, include a directive that tells the model to perform the task step-by-step and to output the result of each step. ```csharp prompt= """Instructions: Compare the pros and cons of EVs and petroleum-fueled vehicles. Break the task into steps, and output the result of each step as you perform it."""; ``` ## Use chain of thought prompting in examples Use examples to indicate the steps for chain of thought prompting, which the model interprets to mean it should also output step results. Steps can include formatting cues. ```csharp prompt= """ Instructions: Compare the pros and cons of EVs and petroleum-fueled vehicles. Differences between EVs and petroleum-fueled vehicles: - Differences ordered according to overall impact, highest-impact first: 1. Summary of vehicle type differences as pros and cons: Pros of EVs 1. Pros of petroleum-fueled vehicles 1. """; ``` ## Related content - [Prompt engineering techniques](/azure/ai-services/openai/concepts/advanced-prompt-engineering) -
data-ingestion.md 12.7 KB
--- title: Data ingestion description: Introduction to data ingestion author: luisquintanilla ms.author: luquinta ms.date: 12/02/2025 ms.topic: concept-article ai-usage: ai-assisted --- # Data ingestion Data ingestion is the process of collecting, reading, and preparing data from different sources such as files, databases, APIs, or cloud services so it can be used in downstream applications. In practice, this process follows the Extract-Transform-Load (ETL) workflow: - **Extract** data from its original source, whether that is a PDF, Word document, audio file, or web API. - **Transform** the data by cleaning, chunking, enriching, or converting formats. - **Load** the data into a destination like a database, vector store, or AI model for retrieval and analysis. For AI and machine learning scenarios, especially Retrieval-Augmented Generation (RAG), data ingestion is not just about converting data from one format to another. It is about making data usable for intelligent applications. This means representing documents in a way that preserves their structure and meaning, splitting them into manageable chunks, enriching them with metadata or embeddings, and storing them so they can be retrieved quickly and accurately. ## Why data ingestion matters for AI applications Imagine you're building a RAG-powered chatbot to help employees find information across your company's vast collection of documents. These documents might include PDFs, Word files, PowerPoint presentations, and web pages scattered across different systems. Your chatbot needs to understand and search through thousands of documents to provide accurate, contextual answers. But raw documents aren't suitable for AI systems. You need to transform them into a format that preserves meaning while making them searchable and retrievable. This is where data ingestion becomes critical. You need to extract text from different file formats, break large documents into smaller chunks that fit within AI model limits, enrich the content with metadata, generate embeddings for semantic search, and store everything in a way that enables fast retrieval. Each step requires careful consideration of how to preserve the original meaning and context. ## The Microsoft.Extensions.DataIngestion library The [📦 Microsoft.Extensions.DataIngestion package](https://www.nuget.org/packages/Microsoft.Extensions.DataIngestion) provides foundational .NET building blocks for data ingestion. It enables developers to read, process, and prepare documents for AI and machine learning workflows, especially Retrieval-Augmented Generation (RAG) scenarios. With these building blocks, you can create robust, flexible, and intelligent data ingestion pipelines tailored for your application needs: - **Unified document representation:** Represent any file type (for example, PDF, Image, or Microsoft Word) in a consistent format that works well with large language models. - **Flexible data ingestion:** Read documents from both cloud services and local sources using multiple built-in readers, making it easy to bring in data from wherever it lives. - **Built-in AI enhancements:** Automatically enrich content with summaries, sentiment analysis, keyword extraction, and classification, preparing your data for intelligent workflows. - **Customizable chunking strategies:** Split documents into chunks using token-based, section-based, or semantic-aware approaches, so you can optimize for your retrieval and analysis needs. - **Production-ready storage:** Store processed chunks in popular vector databases and document stores, with support for embedding generation, making your pipelines ready for real-world scenarios. - **End-to-end pipeline composition:** Chain together readers, processors, chunkers, and writers with the <xref:Microsoft.Extensions.DataIngestion.IngestionPipeline`1> API, reducing boilerplate and making it easy to build, customize, and extend complete workflows. - **Performance and scalability:** Designed for scalable data processing, these components can handle large volumes of data efficiently, making them suitable for enterprise-grade applications. All of these components are open and extensible by design. You can add custom logic and new connectors, and extend the system to support emerging AI scenarios. By standardizing how documents are represented, processed, and stored, .NET developers can build reliable, scalable, and maintainable data pipelines without "reinventing the wheel" for every project. ### Built on stable foundations These data ingestion building blocks are built on top of proven and extensible components in the .NET ecosystem, ensuring reliability, interoperability, and seamless integration with existing AI workflows: - **Microsoft.ML.Tokenizers:** Tokenizers provide the foundation for chunking documents based on tokens. This enables precise splitting of content, which is essential for preparing data for large language models and optimizing retrieval strategies. - **Microsoft.Extensions.AI:** This set of libraries powers enrichment transformations using large language models. It enables features like summarization, sentiment analysis, keyword extraction, and embedding generation, making it easy to enhance your data with intelligent insights. - **Microsoft.Extensions.VectorData:** This set of libraries offers a consistent interface for storing processed chunks in a wide variety of vector stores, including Qdrant, Azure SQL, CosmosDB, MongoDB, ElasticSearch, and many more. This ensures your data pipelines are ready for production and can scale across different storage backends. In addition to familiar patterns and tools, these abstractions build on already extensible components. Plug-in capability and interoperability are paramount, so as the rest of the .NET AI ecosystem grows, the capabilities of the data ingestion components grow as well. This approach empowers developers to easily integrate new connectors, enrichments, and storage options, keeping their pipelines future-ready and adaptable to evolving AI scenarios. ## Data ingestion building blocks The [Microsoft.Extensions.DataIngestion](https://www.nuget.org/packages/Microsoft.Extensions.DataIngestion) library is built around several key components that work together to create a complete data processing pipeline. This section explores each component and how they fit together. ### Documents and document readers At the foundation of the library is the <xref:Microsoft.Extensions.DataIngestion.IngestionDocument> type, which provides a unified way to represent any file format without losing important information. `IngestionDocument` is Markdown-centric because large language models work best with Markdown formatting. The <xref:Microsoft.Extensions.DataIngestion.IngestionDocumentReader> abstraction handles loading documents from various sources, whether local files or streams. A few readers are available: - **[MarkItDown](https://www.nuget.org/packages/Microsoft.Extensions.DataIngestion.MarkItDown)** - **[Markdig](https://www.nuget.org/packages/Microsoft.Extensions.DataIngestion.Markdig/)** More readers (including **LlamaParse** and **Azure Document Intelligence**) will be added in the future. This design means you can work with documents from different sources using the same consistent API, making your code more maintainable and flexible. ### Document processing Document processors apply transformations at the document level to enhance and prepare content. The library provides the <xref:Microsoft.Extensions.DataIngestion.ImageAlternativeTextEnricher> class as a built-in processor that uses large language models to generate descriptive alternative text for images within documents. ### Chunks and chunking strategies Once you have a document loaded, you typically need to break it down into smaller pieces called chunks. Chunks represent subsections of a document that can be efficiently processed, stored, and retrieved by AI systems. This chunking process is essential for retrieval-augmented generation scenarios where you need to find the most relevant pieces of information quickly. The library provides several chunking strategies to fit different use cases: - **Header-based chunking** to split on headers. - **Section-based chunking** to split on sections (for example, pages). - **Semantic-aware chunking** to preserve complete thoughts. These chunking strategies build on the Microsoft.ML.Tokenizers library to intelligently split text into appropriately sized pieces that work well with large language models. The right chunking strategy depends on your document types and how you plan to retrieve information. ```csharp Tokenizer tokenizer = TiktokenTokenizer.CreateForModel("gpt-5"); IngestionChunkerOptions options = new(tokenizer) { MaxTokensPerChunk = 2000, OverlapTokens = 0 }; IngestionChunker<string> chunker = new HeaderChunker(options); ``` ### Chunk processing and enrichment After documents are split into chunks, you can apply processors to enhance and enrich the content. Chunk processors work on individual pieces and can perform: - **Content enrichment** including automatic summaries (`SummaryEnricher`), sentiment analysis (`SentimentEnricher`), and keyword extraction (`KeywordEnricher`). - **Classification** for automated content categorization based on predefined categories (`ClassificationEnricher`). These processors use [Microsoft.Extensions.AI.Abstractions](https://www.nuget.org/packages/Microsoft.Extensions.AI.Abstractions) to leverage large language models for intelligent content transformation, making your chunks more useful for downstream AI applications. ### Document writer and storage <xref:Microsoft.Extensions.DataIngestion.IngestionChunkWriter`1> stores processed chunks into a data store for later retrieval. Using Microsoft.Extensions.AI and [Microsoft.Extensions.VectorData.Abstractions](https://www.nuget.org/packages/Microsoft.Extensions.VectorData.Abstractions), the library provides the <xref:Microsoft.Extensions.DataIngestion.VectorStoreWriter`1> class that supports storing chunks in any vector store supported by Microsoft.Extensions.VectorData. Vector stores include popular options like [Qdrant](https://www.nuget.org/packages/Microsoft.SemanticKernel.Connectors.Qdrant), [SQL Server](https://www.nuget.org/packages/Microsoft.SemanticKernel.Connectors.SqlServer), [CosmosDB](https://www.nuget.org/packages/Microsoft.SemanticKernel.Connectors.CosmosNoSQL), [MongoDB](https://www.nuget.org/packages/Microsoft.SemanticKernel.Connectors.MongoDB), [ElasticSearch](https://www.nuget.org/packages/Elastic.SemanticKernel.Connectors.Elasticsearch), and many more. The writer can also automatically generate embeddings for your chunks using Microsoft.Extensions.AI, readying them for semantic search and retrieval scenarios. ```csharp OpenAIClient openAIClient = new( new ApiKeyCredential(Environment.GetEnvironmentVariable("GITHUB_TOKEN")!), new OpenAIClientOptions { Endpoint = new Uri("https://models.github.ai/inference") }); IEmbeddingGenerator<string, Embedding<float>> embeddingGenerator = openAIClient.GetEmbeddingClient("text-embedding-3-small").AsIEmbeddingGenerator(); using SqliteVectorStore vectorStore = new( "Data Source=vectors.db;Pooling=false", new() { EmbeddingGenerator = embeddingGenerator }); // The writer requires the embedding dimension count to be specified. // For OpenAI's `text-embedding-3-small`, the dimension count is 1536. using VectorStoreWriter<string> writer = new(vectorStore, dimensionCount: 1536); ``` ### Document processing pipeline The <xref:Microsoft.Extensions.DataIngestion.IngestionPipeline`1> API allows you to chain together the various data ingestion components into a complete workflow. You can combine: - **Readers** to load documents from various sources. - **Processors** to transform and enrich document content. - **Chunkers** to break documents into manageable pieces. - **Writers** to store the final results in your chosen data store. This pipeline approach reduces boilerplate code and makes it easy to build, test, and maintain complex data ingestion workflows. ```csharp using IngestionPipeline<string> pipeline = new(reader, chunker, writer, loggerFactory: loggerFactory) { DocumentProcessors = { imageAlternativeTextEnricher }, ChunkProcessors = { summaryEnricher } }; await foreach (var result in pipeline.ProcessAsync(new DirectoryInfo("."), searchPattern: "*.md")) { Console.WriteLine($"Completed processing '{result.DocumentId}'. Succeeded: '{result.Succeeded}'."); } ``` A single document ingestion failure shouldn't fail the whole pipeline. That's why <xref:Microsoft.Extensions.DataIngestion.IngestionPipeline`1.ProcessAsync*?displayProperty=nameWithType> implements partial success by returning `IAsyncEnumerable<IngestionResult>`. The caller is responsible for handling any failures (for example, by retrying failed documents or stopping on first error). -
embeddings.md 5.1 KB
--- title: "How Embeddings Extend Your AI Model's Reach" description: "Learn how embeddings extend the limits and capabilities of AI models in .NET." ms.topic: concept-article #Don't change. ms.date: 03/04/2026 ai-usage: ai-assisted #customer intent: As a .NET developer, I want to understand how embeddings extend LLM limits and capabilities in .NET so that I have more semantic context and better outcomes for my AI apps. --- # Embeddings in .NET Embeddings are the way LLMs capture semantic meaning. They're numeric representations of non-numeric data that an LLM can use to determine relationships between concepts. Use embeddings to help an AI model understand the meaning of inputs so that it can perform comparisons and transformations, such as summarizing text or creating images from text descriptions. LLMs can use embeddings immediately, and you can store embeddings in vector databases to provide semantic memory for LLMs as needed. ## Use cases for embeddings ### Use your own data to improve completion relevance Use your own databases to generate embeddings for your data and integrate it with an LLM to make it available for completions. This use of embeddings is an important component of [retrieval-augmented generation](rag.md). ### Increase the amount of text you can fit in a prompt Use embeddings to increase the amount of context you can fit in a prompt without increasing the number of tokens required. For example, suppose you want to include 500 pages of text in a prompt. The number of tokens for that much raw text exceeds the input token limit, making it impossible to directly include in a prompt. You can use embeddings to summarize and break down large amounts of that text into pieces that are small enough to fit in one input, and then assess the similarity of each piece to the entire raw text. Then you can choose a piece that best preserves the semantic meaning of the raw text and use it in your prompt without hitting the token limit. ### Perform text classification, summarization, or translation Use embeddings to help a model understand the meaning and context of text, and then classify, summarize, or translate that text. For example, you can use embeddings to help models classify texts as positive or negative, spam or not spam, or news or opinion. ### Generate and transcribe audio Use audio embeddings to process audio files or inputs in your app. For example, [Azure Speech in Foundry Tools](/azure/ai-services/speech-service/) supports a range of audio embeddings, including [speech to text](/azure/ai-services/speech-service/speech-to-text) and [text to speech](/azure/ai-services/speech-service/text-to-speech). You can process audio in real-time or in batches. ### Turn text into images or images into text Semantic image processing requires image embeddings, which most LLMs can't generate. Use an image-embedding model such as [ViT](https://huggingface.co/docs/transformers/main/en/model_doc/vit) to create vector embeddings for images. Then you can use those embeddings with an image generation model to create or modify images using text or vice versa. For example, you can [use the DALL·E model to generate images](/azure/ai-services/openai/dall-e-quickstart?tabs=dalle3%2Ccommand-line&pivots=programming-language-csharp) such as logos, faces, animals, and landscapes. ### Generate or document code Use embeddings to help a model create code from text or vice versa, by converting different code or text expressions into a common representation. For example, you can use embeddings to help a model generate or document code in C# or Python. ## Choose an embedding model You generate embeddings for your raw data by using an AI embedding model, which can encode non-numeric data into a vector (a long array of numbers). The model can also decode an embedding into non-numeric data that has the same or similar meaning as the original, raw data. OpenAI's `text-embedding-3-small` and `text-embedding-3-large` are the currently recommended embedding models, replacing the older `text-embedding-ada-002`. For more examples, see the list of [Embedding models available on Azure OpenAI](/azure/ai-services/openai/concepts/models#embeddings). ### Store and process embeddings in a vector database After you generate embeddings, you need a way to store them so you can later retrieve them with calls to an LLM. Vector databases are designed to store and process vectors, so they're a natural home for embeddings. Different vector databases offer different processing capabilities. Choose one based on your raw data and your goals. For information about your options, see [Vector databases for .NET + AI](vector-databases.md). ### Using embeddings in your LLM solution When building LLM-based applications, you can use Agent Framework to integrate embedding models and vector stores, so you can quickly pull in text data, and generate and store embeddings. This lets you use a vector database solution to store and retrieve semantic memories. ## Related content - [How GenAI and LLMs work](how-genai-and-llms-work.md) - [Retrieval-augmented generation](rag.md) - [Training: Develop an AI agent with Microsoft Agent Framework](/training/modules/develop-ai-agent-with-semantic-kernel/) -
how-genai-and-llms-work.md 8.2 KB
--- title: "How Generative AI and LLMs work" description: "Understand how Generative AI and large language models (LLMs) work and how they might be useful in your .NET projects." ms.topic: concept-article ms.date: 03/04/2026 ai-usage: ai-assisted #customer intent: As a .NET developer, I want to understand how Generative AI and large language models (LLMs) work and how they may be useful in my .NET projects. --- # How generative AI and LLMs work Generative AI is a type of artificial intelligence that can create original content, such as natural language, images, audio, and code. The output depends on the inputs you provide. Users commonly interact with generative AI through chat applications that use natural language as input. ChatGPT, developed by OpenAI, is a popular example. Generative AI applications that use natural language as input are powered by large language models (LLMs) to perform natural language processing (NLP). ## How generative AI works All generative AI is built on models. These models are trained with large sets of data in the form of content, such as natural language, images, audio, and code. Generative AI models use the patterns identified in the training data to produce new, statistically similar content. The AI model uses your input to build an output. The model first parses the input into a form it can understand. The model then uses that data to identify matching patterns from its training that it combines to build the final output. Generative AI models are designed to produce unique content, so they won't generate the same output for identical inputs. Generative AI applications that support natural language as input or output use LLMs to do so. The LLM performs NLP, which classifies the input text and determines its sentiment. The generative AI model uses that classification and sentiment analysis to identify patterns and build the output. If the output is text, the LLM alone generates it. If the output is audio or images, additional models provide the data and patterns for generating outputs in that format. ## Common uses of generative AI Generative AI applications support a variety of use cases and outputs, described in the following sections. ### Natural language generation Some generative AI applications can produce natural language outputs, such as: - Professionally worded summaries for a LinkedIn profile - Rich descriptions of items in an online storefront - Ideas for healthy meals For example, given the user input *Give me three ideas for a healthy breakfast including peppers*, a generative AI application might produce the following output: *OK, here are three healthy breakfasts that include peppers:* 1. *Egg and Pepper Scramble: Dice up some bell peppers and sauté them in a pan until they're slightly softened. Then, add in some beaten eggs and scramble everything together. Serve with whole-grain toast for a nutritious breakfast.* 2. *Pepper and Avocado Toast: Toast a slice of whole-grain bread and top it with sliced bell peppers, mashed avocado, and a sprinkle of salt and pepper. This is a great way to get some healthy fats from the avocado along with the vitamin C from the peppers.* 3. *Breakfast Burrito with Peppers: Sauté some diced bell peppers, onions, and scrambled eggs together in a pan. Then, spoon the mixture onto a whole-grain tortilla and top with some salsa or hot sauce. Roll it up and enjoy a satisfying and nutritious breakfast.* ### Image generation Some generative AI applications produce image outputs from natural language inputs. Depending on the application and its models, users might be able to specify: - The people or things to appear in the image - The artistic style to create the image in - References for generating similar images Image generation can create virtual avatars for online accounts, design logos for a business, or provide artistic inspiration for creators. For example, a user might input the request, *Create an image of an elephant eating a burger*. A generative AI application might produce the following output: ### Audio generation Some generative AI applications produce audio outputs from natural language inputs. Depending on the application and its models, users might be able to: - Synthesize natural sounding voices from input text - Create music in a specific style or featuring certain instruments - Modify input audio files based on set criteria provided in natural language Audio generation can provide spoken responses from digital voice assistants, add backing instruments to songs for music production, or reproduce a user's original voice from reference recordings. ### Code generation Some generative AI applications produce code from natural language inputs. Depending on the application and its models, users might be able to generate code that: - Is in the language of their choice - Accepts a specific input and output structure - Uses specific technologies based on their reference documentation Code generation can provide quick implementations of common operations, such as sort or filter data, speed up translating code from one language to another, or answer questions about a technology based on specified reference documentation. For example, given the input "Show me how to code a game of tic-tac-toe with C#", a generative AI application might produce the following output: *Here's a simple example of how you can code a game of tic-tac-toe in C#:* ```csharp using System; namespace TicTacToe { class Program { static char[] board = { '1', '2', '3', '4', '5', '6', '7', '8', '9' }; static int player = 1; // Player 1 will start static int choice; // The number the player chooses static int flag = 0; // Set the flag value to 0 // Check if the position is already marked or not private static int CheckPosition(char mark) { for (int i = 0; i < 9; i++) { if (board[i] == mark) { return 1; } } return 0; } // The rest of the generated code has been omitted for brevity // ... } } ``` *This code creates a simple console-based tic-tac-toe game in C#. It uses a single-dimensional array to represent the board and checks for a win or draw after each move.* ## How LLMs work When training an LLM, the training text is first broken down into [tokens](understanding-tokens.md). Each token identifies a unique text value. A token can be a distinct word, a partial word, or a combination of words and punctuation. Each token is assigned an ID, which enables the text to be represented as a sequence of token IDs. After the text has been broken down into tokens, a contextual vector, known as an [embedding](embeddings.md), is assigned to each token. These embedding vectors are multi-valued numeric data where each element of a token's vector represents a semantic attribute of the token. The elements of a token's vector are determined based on how commonly tokens are used together or in similar contexts. The goal is to predict the next token in the sequence based on the preceding tokens. The model assigns a weight to each token in the existing sequence, representing its relative influence on the next token. The model then uses the preceding tokens' weights and embeddings to calculate and predict the next vector value. The model then selects the most probable token to continue the sequence based on the predicted vector. This process continues iteratively for each token in the sequence, with the output sequence being used regressively as the input for the next iteration. The output is built one token at a time. This strategy is analogous to how auto-complete works, where suggestions are based on what's been typed so far and updated with each new input. During training, the model knows the complete token sequence but ignores all tokens after the one currently being considered. The model compares the predicted vector value to the actual value and calculates the loss. Training then incrementally adjusts the weights to reduce the loss and improve the model. ## Related content - [Understand Tokens](understanding-tokens.md) - [Prompt engineering](prompt-engineering-dotnet.md) - [Large language models](/training/modules/fundamentals-generative-ai/3-language%20models) -
prompt-engineering-dotnet.md 4.7 KB
--- title: Prompt engineering concepts description: Learn basic prompt engineering concepts and how to implement them using .NET tools such as Microsoft Agent Framework. ms.topic: concept-article ms.date: 03/04/2026 ai-usage: ai-assisted --- # Prompt engineering in .NET In this article, you explore essential prompt engineering concepts. Many AI models are prompt-based, meaning they respond to user input text (a *prompt*) with a response generated by predictive algorithms (a *completion*). Newer models also often support completions in chat form, with messages based on roles (system, user, assistant) and chat history to preserve conversations. ## Work with prompts Models that support chat-based apps use three roles to organize completions: a *system* role that controls the chat, a *user* role to represent user input, and an *assistant* role for responding to users. Divide your prompts into messages for each role: - [*System messages*](/azure/ai-services/openai/concepts/advanced-prompt-engineering?pivots=programming-language-chat-completions#system-message) give the model instructions about the assistant. A prompt can have only one system message, and it must be the first message. - *User messages* include prompts from the user, examples, or instructions for the assistant. An example chat completion must have at least one user message. - *Assistant messages* show example or historical completions and must contain a response to the preceding user message. Assistant messages aren't required, but if you include one, it must be paired with a user message to form an example. ## Use instructions to improve the completion An *instruction* is text that tells the model how to respond. An instruction can be a *directive* or an *imperative*: - *Directives* tell the model how to behave but aren't simple commands—think character setup for an improv actor: **"You're helping students learn about U.S. history, so talk about the U.S. unless they specifically ask about other countries or regions."** - *Imperatives* are unambiguous commands for the model to follow. **"Translate to Tagalog:"** ## Use examples to guide the model An example is text that shows the model how to respond by providing sample user input and model output. The model uses examples to infer what to include in completions. Examples can come either before or after the instructions in an engineered prompt, but the two shouldn't be interspersed. An example starts with a prompt and can optionally include a completion. A completion in an example doesn't have to include the verbatim response—it might just contain a formatted word, the first bullet in an unordered list, or something similar to indicate how each completion should start. Classify examples as [zero-shot learning](zero-shot-learning.md#zero-shot-learning) or [few-shot learning](zero-shot-learning.md#few-shot-learning) based on whether they contain verbatim completions. - **Zero-shot learning** examples include a prompt with no verbatim completion. This approach tests a model's responses without giving it example data output. Zero-shot prompts can have completions that include cues, such as indicating the model should output an ordered list by including **"1."** as the completion. - **Few-shot learning** examples include several pairs of prompts with verbatim completions. Few-shot learning can change the model's behavior by adding to its existing knowledge. ## Cues A *cue* is text that conveys the desired structure or format of output. Like an instruction, a cue isn't processed by the model as if it were user input. Like an example, a cue shows the model what you want instead of telling it what to do. Add as many cues as you want to iterate toward the result you want. Use cues with an instruction or an example, and place them at the end of the prompt. ## Example prompt using .NET .NET provides various tools to prompt and chat with different AI models. Use [Agent Framework](/agent-framework/) to connect to a wide variety of AI models and services. Agent Framework includes tools to create agents with system instructions and maintain conversation state across multiple turns. Consider the following code example: The preceding code: - Creates an Azure OpenAI client with an endpoint and API key. - Gets a chat client for the GPT-4o model and converts it to an AI agent. - Creates an agent session to maintain conversation state across multiple turns. - Accepts user input in a loop to allow for different types of prompts. - Asynchronously streams the AI response and displays it to the console. ## Related content - [Prompt engineering techniques](/azure/ai-foundry/openai/concepts/prompt-engineering) - [System message design](/azure/ai-services/openai/concepts/advanced-prompt-engineering) -
rag.md 2.6 KB
--- title: "Integrate Your Data into AI Apps with Retrieval-Augmented Generation" description: "Learn how retrieval-augmented generation lets you use your data with LLMs to generate better completions in .NET." ms.topic: concept-article ms.date: 12/10/2025 --- # Retrieval-augmented generation (RAG) provides LLM knowledge This article describes how retrieval-augmented generation lets LLMs treat your data sources as knowledge without having to train. LLMs have extensive knowledge bases through training. For most scenarios, you can select an LLM that is designed for your requirements, but those LLMs still require additional training to understand your specific data. Retrieval-augmented generation lets you make your data available to LLMs without training them on it first. ## How RAG works To perform retrieval-augmented generation, you create embeddings for your data along with common questions about it. You can do this on the fly or you can create and store the embeddings by using a vector database solution. When a user asks a question, the LLM uses your embeddings to compare the user's question to your data and find the most relevant context. This context and the user's question then go to the LLM in a prompt, and the LLM provides a response based on your data. ### Basic RAG process To perform RAG, you must process each data source that you want to use for retrievals. The basic process is as follows: 1. Chunk large data into manageable pieces. 1. Convert the chunks into a searchable format. 1. Store the converted data in a location that allows efficient access. Additionally, it's important to store relevant metadata for citations or references when the LLM provides responses. 1. Feed your converted data to LLMs in prompts. - **Source data**: This is where your data exists. It could be a file/folder on your machine, a file in cloud storage, an Azure Machine Learning data asset, a Git repository, or an SQL database. - **Data chunking**: The data in your source needs to be converted to plain text. For example, word documents or PDFs need to be cracked open and converted to text. The text is then chunked into smaller pieces. - **Converting the text to vectors**: These are embeddings. Vectors are numerical representations of concepts converted to number sequences, which make it easy for computers to understand the relationships between those concepts. - **Links between source data and embeddings**: This information is stored as metadata on the chunks you created, which are then used to help the LLMs generate citations while generating responses. ## See also - [Data ingestion](data-ingestion.md) -
understanding-tokens.md 6.7 KB
--- title: "Understanding tokens" description: "Understand how large language models (LLMs) use tokens to analyze semantic relationships and generate natural language outputs" ms.topic: concept-article ms.date: 03/04/2026 ai-usage: ai-assisted #customer intent: As a .NET developer, I want understand how large language models (LLMs) use tokens so I can add semantic analysis and text generation capabilities to my .NET projects. --- # Understand tokens When you work with a large language model (LLM), text is first broken into units called *tokens*, which are words, character sets, or combinations of words and punctuation, by a tokenizer. During training, tokenization runs as the first step. The LLM analyzes the semantic relationships between tokens, such as how commonly they're used together or whether they're used in similar contexts. After training, the LLM uses those patterns and relationships to generate a sequence of output tokens based on the input sequence. ## Turn text into tokens The set of unique tokens that an LLM is trained on is known as its _vocabulary_. For example, consider the following sentence: > `I heard a dog bark loudly at a cat` This text could be tokenized as: - `I` - `heard` - `a` - `dog` - `bark` - `loudly` - `at` - `a` - `cat` By having a sufficiently large set of training text, tokenization can compile a vocabulary of many thousands of tokens. ## Common tokenization methods The specific tokenization method varies by LLM. Common tokenization methods include: - **Word** tokenization (text is split into individual words based on a delimiter) - **Character** tokenization (text is split into individual characters) - **Subword** tokenization (text is split into partial words or character sets) For example, the GPT models, developed by OpenAI, use a type of subword tokenization that's known as _Byte-Pair Encoding_ (BPE). OpenAI provides [a tool to visualize how text will be tokenized](https://platform.openai.com/tokenizer). Each tokenization method has benefits and disadvantages: | Token size | Pros | Cons | |----------------------------------------------------|------|------| | Smaller tokens (character or subword tokenization) | - Enables the model to handle a wider range of inputs, such as unknown words, typos, or complex syntax.<br>- Might allow the vocabulary size to be reduced, requiring fewer memory resources. | - A given text is broken into more tokens, requiring additional computational resources while processing.<br>- Given a fixed token limit, the maximum size of the model's input and output is smaller. | | Larger tokens (word tokenization) | - A given text is broken into fewer tokens, requiring fewer computational resources while processing.<br>- Given the same token limit, the maximum size of the model's input and output is larger. | - Might cause an increased vocabulary size, requiring more memory resources.<br>- Can limit the model's ability to handle unknown words, typos, or complex syntax. | ## How LLMs use tokens After the LLM completes tokenization, it assigns an ID to each unique token. Consider this example sentence: > `I heard a dog bark loudly at a cat` After the model uses a word tokenization method, it could assign token IDs as follows: - `I` (1) - `heard` (2) - `a` (3) - `dog` (4) - `bark` (5) - `loudly` (6) - `at` (7) - `a` (the "a" token is already assigned an ID of 3) - `cat` (8) By assigning IDs, text can be represented as a sequence of token IDs. The example sentence would be represented as [1, 2, 3, 4, 5, 6, 7, 3, 8]. The sentence "`I heard a cat`" would be represented as [1, 2, 3, 8]. As training continues, the model adds any new tokens in the training text to its vocabulary and assigns each one an ID. For example: - `meow` (9) - `run` (10) These token ID sequences reveal the semantic relationships between tokens. Multi-valued numeric vectors, known as [embeddings](embeddings.md), represent these relationships. The model assigns an embedding to each token based on how commonly it's used together with, or in similar contexts to, the other tokens. After it's trained, a model can calculate an embedding for text that contains multiple tokens. The model tokenizes the text, then calculates an overall embeddings value based on the learned embeddings of the individual tokens. Use this technique for semantic document searches or to add vector stores to an AI. During output generation, the model predicts a vector value for the next token in the sequence. The model then selects the next token from its vocabulary based on this vector value. In practice, the model calculates multiple vectors by using various elements of the previous tokens' embeddings. The model then evaluates all potential tokens from these vectors and selects the most probable one to continue the sequence. Output generation is an iterative operation. The model appends the predicted token to the sequence so far and uses that as the input for the next iteration, building the final output one token at a time. ### Token limits LLMs have a maximum number of tokens for input and output. This limit is often expressed as a combined maximum _context window_ that covers both input and output tokens together. Taken together, a model's token limit and tokenization method determine the maximum length of text that can be provided as input or generated as output. For example, consider a model that has a maximum context window of 100 tokens. The model processes the example sentences as input text: > `I heard a dog bark loudly at a cat` By using a word-based tokenization method, the input is nine tokens. This leaves 91 **word** tokens available for the output. By using a character-based tokenization method, the input is 34 tokens (including spaces). This leaves only 66 **character** tokens available for the output. ### Token-based pricing and rate limiting Generative AI services often use token-based pricing. The cost of each request depends on the number of input and output tokens. Pricing might differ between input and output. For example, see [Azure OpenAI Service pricing](https://azure.microsoft.com/pricing/details/cognitive-services/openai-service/). Generative AI services also enforce a maximum number of tokens per minute (TPM). These rate limits can vary depending on the service region and LLM. For more information about specific regions, see [Azure OpenAI Service quotas and limits](/azure/ai-services/openai/quotas-limits#regional-quota-limits). ## Related content - [Use Microsoft.ML.Tokenizers for text tokenization](../how-to/use-tokenizers.md) - [How generative AI and LLMs work](how-genai-and-llms-work.md) - [Understand embeddings](embeddings.md) - [Work with vector databases](vector-databases.md) -
vector-databases.md 3.5 KB
--- title: "Using Vector Databases to Extend LLM Capabilities" description: "Learn how vector databases extend LLM capabilities by storing and processing embeddings in .NET." ms.topic: concept-article ms.date: 03/04/2026 ai-usage: ai-assisted --- # Vector databases for .NET + AI Vector databases are designed to store and manage vector [embeddings](embeddings.md). Embeddings are numeric representations of non-numeric data that preserve semantic meaning. You can vectorize words, documents, images, audio, and other data types. Use embeddings to help an AI model understand the meaning of inputs so that it can perform comparisons and transformations, such as summarizing text, finding contextually related data, or creating images from text descriptions. For example, you can use a vector database to: - Identify similar images, documents, and songs based on their contents, themes, sentiments, and styles. - Identify similar products based on their characteristics, features, and user groups. - Recommend content, products, or services based on user preferences. - Identify the best potential options from a large pool of choices to meet complex requirements. - Identify data anomalies or fraudulent activities that are dissimilar from predominant or normal patterns. ## Understand vector search Vector databases provide vector search capabilities to find similar items based on their data characteristics rather than by exact matches on a property field. Vector search works by analyzing the vector representations of your data that you created using an AI embedding model such as the [Azure OpenAI embedding models](/azure/ai-services/openai/concepts/models#embeddings-models). The search process measures the distance between the data vectors and your query vector. The data vectors that are closest to your query vector are the ones that are found to be most similar semantically. Some services such as [Azure Cosmos DB for MongoDB vCore](/azure/cosmos-db/mongodb/vcore/vector-search) provide native vector search capabilities for your data. Other databases can be enhanced with vector search by indexing the stored data using a service such as Azure AI Search, which can scan and index your data to provide vector search capabilities. ## Vector search workflows with .NET and OpenAI Vector databases and their search features are especially useful in [RAG pattern](rag.md) workflows with Azure OpenAI. This pattern lets you augment your AI model with additional semantically rich knowledge of your data. A common AI workflow using vector databases includes these steps: 1. Create embeddings for your data using an OpenAI embedding model. 1. Store and index the embeddings in a vector database or search service. 1. Convert user prompts from your application to embeddings. 1. Run a vector search across your data, comparing the user prompt embedding to the embeddings in your database. 1. Use a language model such as GPT-4o to assemble a user-friendly completion from the vector search results. Visit the [Implement Azure OpenAI with RAG using vector search in a .NET app](../tutorials/tutorial-ai-vector-search.md) tutorial for a hands-on example of this flow. Other benefits of the RAG pattern include: - Generate contextually relevant and accurate responses to user prompts from AI models. - Overcome LLM token limits—the database vector search does the heavy lifting. - Reduce the costs from frequent fine-tuning on updated data. ## Related content - [Implement Azure OpenAI with RAG using vector search in a .NET app](../tutorials/tutorial-ai-vector-search.md) -
zero-shot-learning.md 4.8 KB
--- title: "Zero-shot and few-shot learning" description: "Learn the use cases for zero-shot and few-shot learning in prompt engineering." ms.topic: concept-article #Don't change. ms.date: 03/04/2026 ai-usage: ai-assisted #customer intent: As a .NET developer, I want to understand how zero-shot and few-shot learning techniques can help me improve my prompt engineering. --- # Zero-shot and few-shot learning This article explains zero-shot learning and few-shot learning for prompt engineering in .NET, including their primary use cases. GPT model performance benefits from *prompt engineering*, the practice of providing instructions and examples to a model to refine its output. Zero-shot learning and few-shot learning are techniques you can use when providing examples. ## Zero-shot learning Zero-shot learning is the practice of passing prompts that aren't paired with verbatim completions, although you can include completions that consist of cues. Zero-shot learning relies entirely on the model's existing knowledge to generate responses, which reduces the number of tokens created and can help you control costs. However, zero-shot learning doesn't add to the model's knowledge or context. Here's an example zero-shot prompt that tells the model to evaluate user input to determine which of four possible intents the input represents, and then to preface the response with **"Intent: "**. ```csharp prompt = $""" Instructions: What is the intent of this request? If you don't know the intent, don't guess; instead respond with "Unknown". Choices: SendEmail, SendMessage, CompleteTask, CreateDocument, Unknown. User Input: {request} Intent: """; ``` Zero-shot learning has two primary use cases: - **Work with fine-tuned LLMs** - Because it relies on the model's existing knowledge, zero-shot learning isn't as resource-intensive as few-shot learning, and it works well with LLMs that have already been fine-tuned on instruction datasets. You might be able to rely solely on zero-shot learning and keep costs relatively low. - **Establish performance baselines** - Zero-shot learning can help you simulate how your app performs for actual users. This lets you evaluate various aspects of your model's current performance, such as accuracy or precision. In this case, you typically use zero-shot learning to establish a performance baseline and then experiment with few-shot learning to improve performance. ## Few-shot learning Few-shot learning is the practice of passing prompts paired with verbatim completions (few-shot prompts) to show your model how to respond. Compared to zero-shot learning, this means few-shot learning produces more tokens and causes the model to update its knowledge, which can make few-shot learning more resource-intensive. However, few-shot learning also helps the model produce more relevant responses. ```csharp prompt = $""" Instructions: What is the intent of this request? If you don't know the intent, don't guess; instead respond with "Unknown". Choices: SendEmail, SendMessage, CompleteTask, CreateDocument, Unknown. User Input: Can you send a very quick approval to the marketing team? Intent: SendMessage User Input: Can you send the full update to the marketing team? Intent: SendEmail User Input: {request} Intent: """; ``` Few-shot learning has two primary use cases: - **Tuning an LLM** - Because it can add to the model's knowledge, few-shot learning can improve a model's performance. It also causes the model to create more tokens than zero-shot learning does, which can eventually become prohibitively expensive or even infeasible. However, if your LLM isn't fine-tuned yet, you won't always get good performance with zero-shot prompts, and few-shot learning is warranted. - **Fixing performance issues** - You can use few-shot learning as a follow-up to zero-shot learning. In this case, you use zero-shot learning to establish a performance baseline, and then experiment with few-shot learning based on the zero-shot prompts you used. This lets you add to the model's knowledge after seeing how it currently responds, so you can iterate and improve performance while minimizing the number of tokens you introduce. ### Caveats - Example-based learning doesn't work well for complex reasoning tasks. However, adding instructions can help address this. - Few-shot learning requires creating lengthy prompts. Prompts with a large number of tokens can increase computation and latency. This typically means increased costs. There's also a limit to the length of the prompts. - When you use several examples, the model can learn false patterns, such as "Sentiments are twice as likely to be positive than negative." ## Related content - [Prompt engineering techniques](/azure/ai-services/openai/concepts/advanced-prompt-engineering) - [How GenAI and LLMs work](how-genai-and-llms-work.md)
-
-
evaluation
-
evaluate-ai-response.md 6.5 KB
--- title: Quickstart - Evaluate the quality of a model's response description: Learn how to create an MSTest app to evaluate the AI chat response of a language model. ms.date: 03/18/2025 ms.topic: quickstart --- # Quickstart: Evaluate response quality In this quickstart, you create an MSTest app to evaluate the quality of a chat response from an OpenAI model. The test app uses the [Microsoft.Extensions.AI.Evaluation](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation) libraries. > [!NOTE] > This quickstart demonstrates the simplest usage of the evaluation API. Notably, it doesn't demonstrate use of the [response caching](libraries.md#cached-responses) and [reporting](libraries.md#reporting) functionality, which are important if you're authoring unit tests that run as part of an "offline" evaluation pipeline. The scenario shown in this quickstart is suitable in use cases such as "online" evaluation of AI responses within production code and logging scores to telemetry, where caching and reporting aren't relevant. For a tutorial that demonstrates the caching and reporting functionality, see [Tutorial: Evaluate a model's response with response caching and reporting](evaluate-with-reporting.md) ## Prerequisites - [.NET 8 or a later version](https://dotnet.microsoft.com/download) - [Visual Studio Code](https://code.visualstudio.com/) (optional) ## Configure the AI service To provision an Azure OpenAI service and model using the Azure portal, complete the steps in the [Create and deploy an Azure OpenAI Service resource](/azure/ai-services/openai/how-to/create-resource?pivots=web-portal) article. In the "Deploy a model" step, select the `gpt-5` model. ## Create the test app Complete the following steps to create an MSTest project that connects to an AI model. 1. In a terminal window, navigate to the directory where you want to create your app, and create a new MSTest app with the `dotnet new` command: ```dotnetcli dotnet new mstest -o TestAI ``` 1. Navigate to the `TestAI` directory, and add the necessary packages to your app: ```dotnetcli dotnet add package Azure.AI.OpenAI dotnet add package Azure.Identity dotnet add package Microsoft.Extensions.AI.Abstractions dotnet add package Microsoft.Extensions.AI.Evaluation dotnet add package Microsoft.Extensions.AI.Evaluation.Quality dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` 1. Run the following commands to add [app secrets](/aspnet/core/security/app-secrets) for your Azure OpenAI endpoint and tenant ID: ```bash dotnet user-secrets init dotnet user-secrets set AZURE_OPENAI_ENDPOINT <your-Azure-OpenAI-endpoint> dotnet user-secrets set AZURE_TENANT_ID <your-tenant-ID> ``` (Depending on your environment, the tenant ID might not be needed. In that case, remove it from the code that instantiates the <xref:Azure.Identity.DefaultAzureCredential>.) 1. Open the new app in your editor of choice. ## Add the test app code 1. Rename the *Test1.cs* file to *MyTests.cs*, and then open the file and rename the class to `MyTests`. 1. Add the private <xref:Microsoft.Extensions.AI.Evaluation.ChatConfiguration> and chat message and response members to the `MyTests` class. The `s_messages` field is a list that contains two <xref:Microsoft.Extensions.AI.ChatMessage> objects—one instructs the behavior of the chat bot, and the other is the question from the user. 1. Add the `InitializeAsync` method to the `MyTests` class. This method accomplishes the following tasks: - Sets up the <xref:Microsoft.Extensions.AI.Evaluation.ChatConfiguration>. - Sets the <xref:Microsoft.Extensions.AI.ChatOptions>, including the <xref:Microsoft.Extensions.AI.ChatOptions.Temperature> and the <xref:Microsoft.Extensions.AI.ChatOptions.ResponseFormat>. - Fetches the response to be evaluated by calling <xref:Microsoft.Extensions.AI.IChatClient.GetResponseAsync(System.Collections.Generic.IEnumerable{Microsoft.Extensions.AI.ChatMessage},Microsoft.Extensions.AI.ChatOptions,System.Threading.CancellationToken)>, and stores it in a static variable. 1. Add the `GetAzureOpenAIChatConfiguration` method, which creates the <xref:Microsoft.Extensions.AI.IChatClient> that the evaluator uses to communicate with the model. 1. Add a test method to evaluate the model's response. This method does the following: - Invokes the <xref:Microsoft.Extensions.AI.Evaluation.Quality.CoherenceEvaluator> to evaluate the *coherence* of the response. The <xref:Microsoft.Extensions.AI.Evaluation.IEvaluator.EvaluateAsync(System.Collections.Generic.IEnumerable{Microsoft.Extensions.AI.ChatMessage},Microsoft.Extensions.AI.ChatResponse,Microsoft.Extensions.AI.Evaluation.ChatConfiguration,System.Collections.Generic.IEnumerable{Microsoft.Extensions.AI.Evaluation.EvaluationContext},System.Threading.CancellationToken)> method returns an <xref:Microsoft.Extensions.AI.Evaluation.EvaluationResult> that contains a <xref:Microsoft.Extensions.AI.Evaluation.NumericMetric>. A `NumericMetric` contains a numeric value that's typically used to represent numeric scores that fall within a well-defined range. - Retrieves the coherence score from the <xref:Microsoft.Extensions.AI.Evaluation.EvaluationResult>. - Validates the *default interpretation* for the returned coherence metric. Evaluators can include a default interpretation for the metrics they return. You can also change the default interpretation to suit your specific requirements, if needed. - Validates that no diagnostics are present on the returned coherence metric. Evaluators can include diagnostics on the metrics they return to indicate errors, warnings, or other exceptional conditions encountered during evaluation. ## Run the test/evaluation Run the test using your preferred test workflow, for example, by using the CLI command `dotnet test` or through [Test Explorer](/visualstudio/test/run-unit-tests-with-test-explorer). ## Clean up resources If you no longer need them, delete the Azure OpenAI resource and GPT-4 model deployment. 1. In the [Azure portal](https://aka.ms/azureportal), navigate to the Azure OpenAI resource. 1. Select the Azure OpenAI resource, and then select **Delete**. ## Next steps - Evaluate the responses from different OpenAI models. - Add response caching and reporting to your evaluation code. For more information, see [Tutorial: Evaluate a model's response with response caching and reporting](evaluate-with-reporting.md). -
evaluate-safety.md 12.5 KB
--- title: Tutorial - Evaluate response safety with caching and reporting description: Create an MSTest app that evaluates the content safety of a model's response using the evaluators in the Microsoft.Extensions.AI.Evaluation.Safety package and with caching and reporting. ms.date: 05/12/2025 ms.topic: tutorial --- # Tutorial: Evaluate response safety with caching and reporting In this tutorial, you create an MSTest app to evaluate the *content safety* of a response from an OpenAI model. Safety evaluators check for presence of harmful, inappropriate, or unsafe content in a response. The test app uses the safety evaluators from the [Microsoft.Extensions.AI.Evaluation.Safety](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Safety) package to perform the evaluations. These safety evaluators use the [Microsoft Foundry](/azure/ai-foundry/) Evaluation service to perform evaluations. ## Prerequisites - .NET 8.0 SDK or higher - [Install the .NET 8 SDK](https://dotnet.microsoft.com/download/dotnet/8.0). - An Azure subscription - [Create one for free](https://azure.microsoft.com/free). ## Configure the AI service To provision an Azure OpenAI service and model using the Azure portal, complete the steps in the [Create and deploy an Azure OpenAI Service resource](/azure/ai-services/openai/how-to/create-resource?pivots=web-portal) article. In the "Deploy a model" step, select the `gpt-5` model. > [!TIP] > The previous configuration step is only required to fetch the response to be evaluated. To evaluate the safety of a response you already have in hand, you can skip this configuration. The evaluators in this tutorial use the Foundry Evaluation service, which requires some additional setup: - [Create a resource group](/azure/azure-resource-manager/management/manage-resource-groups-portal#create-resource-groups) within one of the Azure [regions that support Foundry Evaluation service](/azure/ai-foundry/how-to/develop/evaluate-sdk#region-support). - [Create a Foundry hub](/azure/ai-foundry/how-to/create-azure-ai-resource?tabs=portal#create-a-hub-in-azure-ai-foundry-portal) in the resource group you just created. - Finally, [create a Foundry project](/azure/ai-foundry/how-to/create-projects?tabs=ai-studio#create-a-project) in the hub you just created. ## Create the test app Complete the following steps to create an MSTest project. 1. In a terminal window, navigate to the directory where you want to create your app, and create a new MSTest app with the `dotnet new` command: ```dotnetcli dotnet new mstest -o EvaluateResponseSafety ``` 1. Navigate to the `EvaluateResponseSafety` directory, and add the necessary packages to your app: ```dotnetcli dotnet add package Azure.AI.OpenAI dotnet add package Azure.Identity dotnet add package Microsoft.Extensions.AI.Abstractions dotnet add package Microsoft.Extensions.AI.Evaluation dotnet add package Microsoft.Extensions.AI.Evaluation.Reporting dotnet add package Microsoft.Extensions.AI.Evaluation.Safety --prerelease dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` 1. Run the following commands to add [app secrets](/aspnet/core/security/app-secrets) for your Azure OpenAI endpoint, tenant ID, subscription ID, resource group, and project: ```bash dotnet user-secrets init dotnet user-secrets set AZURE_OPENAI_ENDPOINT <your-Azure-OpenAI-endpoint> dotnet user-secrets set AZURE_TENANT_ID <your-tenant-ID> dotnet user-secrets set AZURE_SUBSCRIPTION_ID <your-subscription-ID> dotnet user-secrets set AZURE_RESOURCE_GROUP <your-resource-group> dotnet user-secrets set AZURE_AI_PROJECT <your-Azure-AI-project> ``` (Depending on your environment, the tenant ID might not be needed. In that case, remove it from the code that instantiates the <xref:Azure.Identity.DefaultAzureCredential>.) 1. Open the new app in your editor of choice. ## Add the test app code 1. Rename the `Test1.cs` file to `MyTests.cs`, and then open the file and rename the class to `MyTests`. Delete the empty `TestMethod1` method. 1. Add the necessary `using` directives to the top of the file. 1. Add the <xref:Microsoft.VisualStudio.TestTools.UnitTesting.TestContext> property to the class. 1. Add the scenario and execution name fields to the class. The [scenario name](xref:Microsoft.Extensions.AI.Evaluation.Reporting.ScenarioRun.ScenarioName) is set to the fully qualified name of the current test method. However, you can set it to any string of your choice. Here are some considerations for choosing a scenario name: - When using disk-based storage, the scenario name is used as the name of the folder under which the corresponding evaluation results are stored. - By default, the generated evaluation report splits scenario names on `.` so that the results can be displayed in a hierarchical view with appropriate grouping, nesting, and aggregation. The [execution name](xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration.ExecutionName) is used to group evaluation results that are part of the same evaluation run (or test run) when the evaluation results are stored. If you don't provide an execution name when creating a <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration>, all evaluation runs will use the same default execution name of `Default`. In this case, results from one run will be overwritten by the next. 1. Add a method to gather the safety evaluators to use in the evaluation. 1. Add a <xref:Microsoft.Extensions.AI.Evaluation.Safety.ContentSafetyServiceConfiguration> object, which configures the connection parameters that the safety evaluators need to communicate with the Foundry Evaluation service. 1. Add a method that creates an <xref:Microsoft.Extensions.AI.IChatClient> object, which will be used to get the chat response to evaluate from the LLM. 1. Set up the reporting functionality. Convert the <xref:Microsoft.Extensions.AI.Evaluation.Safety.ContentSafetyServiceConfiguration> to a <xref:Microsoft.Extensions.AI.Evaluation.ChatConfiguration>, and then pass that to the method that creates a <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration>. Response caching functionality is supported and works the same way regardless of whether the evaluators talk to an LLM or to the Foundry Evaluation service. The response will be reused until the corresponding cache entry expires (in 14 days by default), or until any request parameter, such as the LLM endpoint or the question being asked, is changed. > [!NOTE] > This code example passes the LLM <xref:Microsoft.Extensions.AI.IChatClient> as `originalChatClient` to <xref:Microsoft.Extensions.AI.Evaluation.Safety.ContentSafetyServiceConfigurationExtensions.ToChatConfiguration(Microsoft.Extensions.AI.Evaluation.Safety.ContentSafetyServiceConfiguration,Microsoft.Extensions.AI.IChatClient)>. The reason to include the LLM chat client here is to enable getting a chat response from the LLM, and notably, to enable response caching for it. (If you don't want to cache the LLM's response, you can create a separate, local <xref:Microsoft.Extensions.AI.IChatClient> to fetch the response from the LLM.) Instead of passing a <xref:Microsoft.Extensions.AI.IChatClient>, if you already have a <xref:Microsoft.Extensions.AI.Evaluation.ChatConfiguration> for an LLM from another reporting configuration, you can pass that instead, using the <xref:Microsoft.Extensions.AI.Evaluation.Safety.ContentSafetyServiceConfigurationExtensions.ToChatConfiguration(Microsoft.Extensions.AI.Evaluation.Safety.ContentSafetyServiceConfiguration,Microsoft.Extensions.AI.Evaluation.ChatConfiguration)> overload. > > Similarly, if you configure both [LLM-based evaluators](libraries.md#quality-evaluators) and [Foundry Evaluation service–based evaluators](libraries.md#safety-evaluators) in the reporting configuration, you also need to pass the LLM <xref:Microsoft.Extensions.AI.Evaluation.ChatConfiguration> to <xref:Microsoft.Extensions.AI.Evaluation.Safety.ContentSafetyServiceConfigurationExtensions.ToChatConfiguration(Microsoft.Extensions.AI.Evaluation.Safety.ContentSafetyServiceConfiguration,Microsoft.Extensions.AI.Evaluation.ChatConfiguration)>. Then it returns a <xref:Microsoft.Extensions.AI.Evaluation.ChatConfiguration> that can talk to both types of evaluators. 1. Add a method to define the [chat options](xref:Microsoft.Extensions.AI.ChatOptions) and ask the model for a response to a given question. The test in this tutorial evaluates the LLM's response to an astronomy question. Since the <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration> has response caching enabled, and since the supplied <xref:Microsoft.Extensions.AI.IChatClient> is always fetched from the <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ScenarioRun> created using this reporting configuration, the LLM response for the test is cached and reused. 1. Add a method to validate the response. > [!TIP] > Some of the evaluators, for example, <xref:Microsoft.Extensions.AI.Evaluation.Safety.ViolenceEvaluator>, might produce a warning diagnostic that's shown [in the report](#generate-a-report) if you only evaluate the response and not the message. Similarly, if the data you pass to <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ScenarioRunExtensions.EvaluateAsync*> contains two consecutive messages with the same <xref:Microsoft.Extensions.AI.ChatRole> (for example, <xref:Microsoft.Extensions.AI.ChatRole.User> or <xref:Microsoft.Extensions.AI.ChatRole.Assistant>), it might also produce a warning. However, even though an evaluator might produce a warning diagnostic in these cases, it still proceeds with the evaluation. 1. Finally, add the [test method](xref:Microsoft.VisualStudio.TestTools.UnitTesting.TestMethodAttribute) itself. This test method: - Creates the <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ScenarioRun>. The use of `await using` ensures that the `ScenarioRun` is correctly disposed and that the results of this evaluation are correctly persisted to the result store. - Gets the LLM's response to a specific astronomy question. The same <xref:Microsoft.Extensions.AI.IChatClient> that will be used for evaluation is passed to the `GetAstronomyConversationAsync` method in order to get *response caching* for the primary LLM response being evaluated. (In addition, this enables response caching for the responses that the evaluators fetch from the Foundry Evaluation service as part of performing their evaluations.) - Runs the evaluators against the response. Like the LLM response, on subsequent runs, the evaluation is fetched from the (disk-based) response cache that was configured in `s_safetyReportingConfig`. - Runs some safety validation on the evaluation result. ## Run the test/evaluation Run the test using your preferred test workflow, for example, by using the CLI command `dotnet test` or through [Test Explorer](/visualstudio/test/run-unit-tests-with-test-explorer). ## Generate a report To generate a report to view the evaluation results, see [Generate a report](evaluate-with-reporting.md#generate-a-report). ## Next steps This tutorial covers the basics of evaluating content safety. As you create your test suite, consider the following next steps: - Configure additional evaluators, such as the [quality evaluators](libraries.md#quality-evaluators). For an example, see the AI samples repo [quality and safety evaluation example](https://github.com/dotnet/ai-samples/blob/main/src/microsoft-extensions-ai-evaluation/api/reporting/ReportingExamples.Example10_RunningQualityAndSafetyEvaluatorsTogether.cs). - Evaluate the content safety of generated images. For an example, see the AI samples repo [image response example](https://github.com/dotnet/ai-samples/blob/main/src/microsoft-extensions-ai-evaluation/api/reporting/ReportingExamples.Example09_RunningSafetyEvaluatorsAgainstResponsesWithImages.cs). - In real-world evaluations, you might not want to validate individual results, since the LLM responses and evaluation scores can vary over time as your product (and the models used) evolve. You might not want individual evaluation tests to fail and block builds in your CI/CD pipelines when this happens. Instead, in such cases, it might be better to rely on the generated report and track the overall trends for evaluation scores across different scenarios over time (and only fail individual builds in your CI/CD pipelines when there's a significant drop in evaluation scores across multiple different tests). -
evaluate-with-reporting.md 14.2 KB
--- title: Tutorial - Evaluate response quality with caching and reporting description: Create an MSTest app to evaluate the response quality of a language model, add a custom evaluator, and learn how to use the caching and reporting features of Microsoft.Extensions.AI.Evaluation. ms.date: 05/09/2025 ms.topic: tutorial --- # Tutorial: Evaluate response quality with caching and reporting In this tutorial, you create an MSTest app to evaluate the chat response of an OpenAI model. The test app uses the [Microsoft.Extensions.AI.Evaluation](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation) libraries to perform the evaluations, cache the model responses, and create reports. The tutorial uses both built-in and custom evaluators. The built-in quality evaluators (from the [Microsoft.Extensions.AI.Evaluation.Quality package](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Quality)) use an LLM to perform evaluations; the custom evaluator does not use AI. ## Prerequisites - [.NET 8 or a later version](https://dotnet.microsoft.com/download) - [Visual Studio Code](https://code.visualstudio.com/) (optional) ## Configure the AI service To provision an Azure OpenAI service and model using the Azure portal, complete the steps in the [Create and deploy an Azure OpenAI Service resource](/azure/ai-services/openai/how-to/create-resource?pivots=web-portal) article. In the "Deploy a model" step, select the `gpt-5` model. ## Create the test app Complete the following steps to create an MSTest project that connects to an AI model. 1. In a terminal window, navigate to the directory where you want to create your app, and create a new MSTest app with the `dotnet new` command: ```dotnetcli dotnet new mstest -o TestAIWithReporting ``` 1. Navigate to the `TestAIWithReporting` directory, and add the necessary packages to your app: ```dotnetcli dotnet add package Azure.AI.OpenAI dotnet add package Azure.Identity dotnet add package Microsoft.Extensions.AI.Abstractions dotnet add package Microsoft.Extensions.AI.Evaluation dotnet add package Microsoft.Extensions.AI.Evaluation.Quality dotnet add package Microsoft.Extensions.AI.Evaluation.Reporting dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` 1. Run the following commands to add [app secrets](/aspnet/core/security/app-secrets) for your Azure OpenAI endpoint and tenant ID: ```bash dotnet user-secrets init dotnet user-secrets set AZURE_OPENAI_ENDPOINT <your-Azure-OpenAI-endpoint> dotnet user-secrets set AZURE_TENANT_ID <your-tenant-ID> ``` (Depending on your environment, the tenant ID might not be needed. In that case, remove it from the code that instantiates the <xref:Azure.Identity.DefaultAzureCredential>.) 1. Open the new app in your editor of choice. ## Add the test app code 1. Rename the *Test1.cs* file to *MyTests.cs*, and then open the file and rename the class to `MyTests`. Delete the empty `TestMethod1` method. 1. Add the necessary `using` directives to the top of the file. 1. Add the <xref:Microsoft.VisualStudio.TestTools.UnitTesting.TestContext> property to the class. 1. Add the `GetAzureOpenAIChatConfiguration` method, which creates the <xref:Microsoft.Extensions.AI.IChatClient> that the evaluator uses to communicate with the model. 1. Set up the reporting functionality. **Scenario name** The [scenario name](xref:Microsoft.Extensions.AI.Evaluation.Reporting.ScenarioRun.ScenarioName) is set to the fully qualified name of the current test method. However, you can set it to any string of your choice when you call <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration.CreateScenarioRunAsync(System.String,System.String,System.Collections.Generic.IEnumerable{System.String},System.Collections.Generic.IEnumerable{System.String},System.Threading.CancellationToken)>. Here are some considerations for choosing a scenario name: - When using disk-based storage, the scenario name is used as the name of the folder under which the corresponding evaluation results are stored. So it's a good idea to keep the name reasonably short and avoid any characters that aren't allowed in file and directory names. - By default, the generated evaluation report splits scenario names on `.` so that the results can be displayed in a hierarchical view with appropriate grouping, nesting, and aggregation. This is especially useful in cases where the scenario name is set to the fully qualified name of the corresponding test method, since it allows the results to be grouped by namespaces and class names in the hierarchy. However, you can also take advantage of this feature by including periods (`.`) in your own custom scenario names to create a reporting hierarchy that works best for your scenarios. **Execution name** The execution name is used to group evaluation results that are part of the same evaluation run (or test run) when the evaluation results are stored. If you don't provide an execution name when creating a <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration>, all evaluation runs will use the same default execution name of `Default`. In this case, results from one run will be overwritten by the next and you lose the ability to compare results across different runs. This example uses a timestamp as the execution name. If you have more than one test in your project, ensure that results are grouped correctly by using the same execution name in all reporting configurations used across the tests. In a more real-world scenario, you might also want to share the same execution name across evaluation tests that live in multiple different assemblies and that are executed in different test processes. In such cases, you could use a script to update an environment variable with an appropriate execution name (such as the current build number assigned by your CI/CD system) before running the tests. Or, if your build system produces monotonically increasing assembly file versions, you could read the <xref:System.Reflection.AssemblyFileVersionAttribute> from within the test code and use that as the execution name to compare results across different product versions. **Reporting configuration** A <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration> identifies: - The set of evaluators that should be invoked for each <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ScenarioRun> that's created by calling <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration.CreateScenarioRunAsync(System.String,System.String,System.Collections.Generic.IEnumerable{System.String},System.Collections.Generic.IEnumerable{System.String},System.Threading.CancellationToken)>. - The LLM endpoint that the evaluators should use (see <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration.ChatConfiguration?displayProperty=nameWithType>). - How and where the results for the scenario runs should be stored. - How LLM responses related to the scenario runs should be cached. - The execution name that should be used when reporting results for the scenario runs. This test uses a disk-based reporting configuration. 1. In a separate file, add the `WordCountEvaluator` class, which is a custom evaluator that implements <xref:Microsoft.Extensions.AI.Evaluation.IEvaluator>. The `WordCountEvaluator` counts the number of words present in the response. Unlike some evaluators, it isn't based on AI. The `EvaluateAsync` method returns an <xref:Microsoft.Extensions.AI.Evaluation.EvaluationResult> includes a <xref:Microsoft.Extensions.AI.Evaluation.NumericMetric> that contains the word count. The `EvaluateAsync` method also attaches a default interpretation to the metric. The default interpretation considers the metric to be good (acceptable) if the detected word count is between 6 and 100. Otherwise, the metric is considered failed. This default interpretation can be overridden by the caller, if needed. 1. Back in `MyTests.cs`, add a method to gather the evaluators to use in the evaluation. 1. Add a method to add a system prompt <xref:Microsoft.Extensions.AI.ChatMessage>, define the [chat options](xref:Microsoft.Extensions.AI.ChatOptions), and ask the model for a response to a given question. The test in this tutorial evaluates the LLM's response to an astronomy question. Since the <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration> has response caching enabled, and since the supplied <xref:Microsoft.Extensions.AI.IChatClient> is always fetched from the <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ScenarioRun> created using this reporting configuration, the LLM response for the test is cached and reused. The response will be reused until the corresponding cache entry expires (in 14 days by default), or until any request parameter, such as the the LLM endpoint or the question being asked, is changed. 1. Add a method to validate the response. > [!TIP] > The metrics each include a `Reason` property that explains the reasoning for the score. The reason is included in the [generated report](#generate-a-report) and can be viewed by clicking on the information icon on the corresponding metric's card. 1. Finally, add the [test method](xref:Microsoft.VisualStudio.TestTools.UnitTesting.TestMethodAttribute) itself. This test method: - Creates the <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ScenarioRun>. The use of `await using` ensures that the `ScenarioRun` is correctly disposed and that the results of this evaluation are correctly persisted to the result store. - Gets the LLM's response to a specific astronomy question. The same <xref:Microsoft.Extensions.AI.IChatClient> that will be used for evaluation is passed to the `GetAstronomyConversationAsync` method in order to get *response caching* for the primary LLM response being evaluated. (In addition, this enables response caching for the LLM turns that the evaluators use to perform their evaluations internally.) With response caching, the LLM response is fetched either: - Directly from the LLM endpoint in the first run of the current test, or in subsequent runs if the cached entry has expired (14 days, by default). - From the (disk-based) response cache that was configured in `s_defaultReportingConfiguration` in subsequent runs of the test. - Runs the evaluators against the response. Like the LLM response, on subsequent runs, the evaluation is fetched from the (disk-based) response cache that was configured in `s_defaultReportingConfiguration`. - Runs some basic validation on the evaluation result. This step is optional and mainly for demonstration purposes. In real-world evaluations, you might not want to validate individual results since the LLM responses and evaluation scores can change over time as your product (and the models used) evolve. You might not want individual evaluation tests to "fail" and block builds in your CI/CD pipelines when this happens. Instead, it might be better to rely on the generated report and track the overall trends for evaluation scores across different scenarios over time (and only fail individual builds when there's a significant drop in evaluation scores across multiple different tests). That said, there is some nuance here and the choice of whether to validate individual results or not can vary depending on the specific use case. When the method returns, the `scenarioRun` object is disposed and the evaluation result for the evaluation is stored to the (disk-based) result store that's configured in `s_defaultReportingConfiguration`. ## Run the test/evaluation Run the test using your preferred test workflow, for example, by using the CLI command `dotnet test` or through [Test Explorer](/visualstudio/test/run-unit-tests-with-test-explorer). ## Generate a report 1. Install the [Microsoft.Extensions.AI.Evaluation.Console](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Console) .NET tool by running the following command from a terminal window: ```dotnetcli dotnet tool install --local Microsoft.Extensions.AI.Evaluation.Console ``` > [!TIP] > You might need to create a manifest file first. For more information about that and installing local tools, see [Local tools](../../core/tools/dotnet-tool-install.md#local-tools). 1. Generate a report by running the following command: ```dotnetcli dotnet tool run aieval report --path <path\to\your\cache\storage> --output report.html ``` 1. Open the `report.html` file. It should look something like this. ## Next steps - Navigate to the directory where the test results are stored (which is `C:\TestReports`, unless you modified the location when you created the <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration>). In the `results` subdirectory, notice that there's a folder for each test run named with a timestamp (`ExecutionName`). Inside each of those folders is a folder for each scenario name—in this case, just the single test method in the project. That folder contains a JSON file with the all the data including the messages, response, and evaluation result. - Expand the evaluation. Here are a couple ideas: - Add an additional custom evaluator, such as [an evaluator that uses AI to determine the measurement system](https://github.com/dotnet/ai-samples/blob/main/src/microsoft-extensions-ai-evaluation/api/evaluation/Evaluators/MeasurementSystemEvaluator.cs) that's used in the response. - Add another test method, for example, [a method that evaluates multiple responses](https://github.com/dotnet/ai-samples/blob/main/src/microsoft-extensions-ai-evaluation/api/reporting/ReportingExamples.Example02_SamplingAndEvaluatingMultipleResponses.cs) from the LLM. Since each response can be different, it's good to sample and evaluate at least a few responses to a question. In this case, you specify an iteration name each time you call <xref:Microsoft.Extensions.AI.Evaluation.Reporting.ReportingConfiguration.CreateScenarioRunAsync(System.String,System.String,System.Collections.Generic.IEnumerable{System.String},System.Collections.Generic.IEnumerable{System.String},System.Threading.CancellationToken)>. -
libraries.md 11.9 KB
--- title: The Microsoft.Extensions.AI.Evaluation libraries description: Learn about the Microsoft.Extensions.AI.Evaluation libraries, which simplify the process of evaluating the quality and accuracy of responses generated by AI models in .NET intelligent apps. ms.topic: concept-article ms.date: 07/24/2025 --- # The Microsoft.Extensions.AI.Evaluation libraries The Microsoft.Extensions.AI.Evaluation libraries simplify the process of evaluating the quality and safety of responses generated by AI models in .NET intelligent apps. Various quality metrics measure aspects like relevance, truthfulness, coherence, and completeness of the responses. Safety metrics measure aspects like hate and unfairness, violence, and sexual content. Evaluations are crucial in testing, because they help ensure that the AI model performs as expected and provides reliable and accurate results. The evaluation libraries, which are built on top of the [Microsoft.Extensions.AI abstractions](../microsoft-extensions-ai.md), are composed of the following NuGet packages: - [📦 Microsoft.Extensions.AI.Evaluation](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation) – Defines the core abstractions and types for supporting evaluation. - [📦 Microsoft.Extensions.AI.Evaluation.NLP](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.NLP) - Contains [evaluators](#nlp-evaluators) that evaluate the similarity of an LLM's response text to one or more reference responses using natural language processing (NLP) metrics. These evaluators aren't LLM or AI-based; they use traditional NLP techniques such as text tokenization and n-gram analysis to evaluate text similarity. - [📦 Microsoft.Extensions.AI.Evaluation.Quality](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Quality) – Contains [evaluators](#quality-evaluators) that assess the quality of LLM responses in an app according to metrics such as relevance and completeness. These evaluators use the LLM directly to perform evaluations. - [📦 Microsoft.Extensions.AI.Evaluation.Safety](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Safety) – Contains [evaluators](#safety-evaluators), such as the `ProtectedMaterialEvaluator` and `ContentHarmEvaluator`, that use the [Microsoft Foundry](/azure/ai-foundry/) Evaluation service to perform evaluations. - [📦 Microsoft.Extensions.AI.Evaluation.Reporting](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Reporting) – Contains support for caching LLM responses, storing the results of evaluations, and generating reports from that data. - [📦 Microsoft.Extensions.AI.Evaluation.Reporting.Azure](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Reporting.Azure) - Supports the reporting library with an implementation for caching LLM responses and storing the evaluation results in an [Azure Storage](/azure/storage/common/storage-introduction) container. - [📦 Microsoft.Extensions.AI.Evaluation.Console](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Console) – A command-line tool for generating reports and managing evaluation data. ## Test integration The libraries are designed to integrate smoothly with existing .NET apps, allowing you to leverage existing testing infrastructures and familiar syntax to evaluate intelligent apps. You can use any test framework (for example, [MSTest](../../core/testing/index.md#mstest), [xUnit](../../core/testing/index.md#xunitnet), or [NUnit](../../core/testing/index.md#nunit)) and testing workflow (for example, [Test Explorer](/visualstudio/test/run-unit-tests-with-test-explorer), [dotnet test](../../core/tools/dotnet-test.md), or a CI/CD pipeline). The library also provides easy ways to do online evaluations of your application by publishing evaluation scores to telemetry and monitoring dashboards. ## Comprehensive evaluation metrics The evaluation libraries were built in collaboration with data science researchers from Microsoft and GitHub, and were tested on popular Microsoft Copilot experiences. The following sections show the built-in [quality](#quality-evaluators), [NLP](#nlp-evaluators), and [safety](#safety-evaluators) evaluators and the metrics they measure. You can also customize to add your own evaluations by implementing the <xref:Microsoft.Extensions.AI.Evaluation.IEvaluator> interface. ### Quality evaluators Quality evaluators measure response quality. They use an LLM to perform the evaluation. | Evaluator type | Metric | Description | |----------------------------------------------------------------------|-------------|-------------| | <xref:Microsoft.Extensions.AI.Evaluation.Quality.RelevanceEvaluator> | `Relevance` | Evaluates how relevant a response is to a query | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.CompletenessEvaluator> | `Completeness` | Evaluates how comprehensive and accurate a response is | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.RetrievalEvaluator> | `Retrieval` | Evaluates performance in retrieving information for additional context | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.FluencyEvaluator> | `Fluency` | Evaluates grammatical accuracy, vocabulary range, sentence complexity, and overall readability| | <xref:Microsoft.Extensions.AI.Evaluation.Quality.CoherenceEvaluator> | `Coherence` | Evaluates the logical and orderly presentation of ideas | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.EquivalenceEvaluator> | `Equivalence` | Evaluates the similarity between the generated text and its ground truth with respect to a query | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.GroundednessEvaluator> | `Groundedness` | Evaluates how well a generated response aligns with the given context | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.RelevanceTruthAndCompletenessEvaluator>† | `Relevance (RTC)`, `Truth (RTC)`, and `Completeness (RTC)` | Evaluates how relevant, truthful, and complete a response is | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.IntentResolutionEvaluator> | `Intent Resolution` | Evaluates an AI system's effectiveness at identifying and resolving user intent (agent-focused) | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.TaskAdherenceEvaluator> | `Task Adherence` | Evaluates an AI system's effectiveness at adhering to the task assigned to it (agent-focused) | | <xref:Microsoft.Extensions.AI.Evaluation.Quality.ToolCallAccuracyEvaluator> | `Tool Call Accuracy` | Evaluates an AI system's effectiveness at using the tools supplied to it (agent-focused) | † This evaluator is marked [experimental](../../fundamentals/syslib-diagnostics/experimental-overview.md). ### NLP evaluators NLP evaluators evaluate the quality of an LLM response by comparing it to a reference response using natural language processing (NLP) techniques. These evaluators aren't LLM or AI-based; instead, they use older NLP techniques to perform text comparisons. | Evaluator type | Metric | Description | |---------------------------------------------------------------------------|--------------------|-------------| | <xref:Microsoft.Extensions.AI.Evaluation.NLP.BLEUEvaluator> | `BLEU` | Evaluates a response by comparing it to one or more reference responses using the bilingual evaluation understudy (BLEU) algorithm. This algorithm is commonly used to evaluate the quality of machine-translation or text-generation tasks. | | <xref:Microsoft.Extensions.AI.Evaluation.NLP.GLEUEvaluator> | `GLEU` | Measures the similarity between the generated response and one or more reference responses using the Google BLEU (GLEU) algorithm, a variant of the BLEU algorithm that's optimized for sentence-level evaluation. | | <xref:Microsoft.Extensions.AI.Evaluation.NLP.F1Evaluator> | `F1` | Evaluates a response by comparing it to a reference response using the *F1* scoring algorithm (the ratio of the number of shared words between the generated response and the reference response). | ### Safety evaluators Safety evaluators check for presence of harmful, inappropriate, or unsafe content in a response. They rely on the Foundry Evaluation service, which uses a model that's fine tuned to perform evaluations. | Evaluator type | Metric | Description | |---------------------------------------------------------------------------|--------------------|-------------| | <xref:Microsoft.Extensions.AI.Evaluation.Safety.GroundednessProEvaluator> | `Groundedness Pro` | Uses a fine-tuned model hosted behind the Foundry Evaluation service to evaluate how well a generated response aligns with the given context | | <xref:Microsoft.Extensions.AI.Evaluation.Safety.ProtectedMaterialEvaluator> | `Protected Material` | Evaluates response for the presence of protected material | | <xref:Microsoft.Extensions.AI.Evaluation.Safety.UngroundedAttributesEvaluator> | `Ungrounded Attributes` | Evaluates a response for the presence of content that indicates ungrounded inference of human attributes | | <xref:Microsoft.Extensions.AI.Evaluation.Safety.HateAndUnfairnessEvaluator>† | `Hate And Unfairness` | Evaluates a response for the presence of content that's hateful or unfair | | <xref:Microsoft.Extensions.AI.Evaluation.Safety.SelfHarmEvaluator>† | `Self Harm` | Evaluates a response for the presence of content that indicates self harm | | <xref:Microsoft.Extensions.AI.Evaluation.Safety.ViolenceEvaluator>† | `Violence` | Evaluates a response for the presence of violent content | | <xref:Microsoft.Extensions.AI.Evaluation.Safety.SexualEvaluator>† | `Sexual` | Evaluates a response for the presence of sexual content | | <xref:Microsoft.Extensions.AI.Evaluation.Safety.CodeVulnerabilityEvaluator> | `Code Vulnerability` | Evaluates a response for the presence of vulnerable code | | <xref:Microsoft.Extensions.AI.Evaluation.Safety.IndirectAttackEvaluator> | `Indirect Attack` | Evaluates a response for the presence of indirect attacks, such as manipulated content, intrusion, and information gathering | † In addition, the <xref:Microsoft.Extensions.AI.Evaluation.Safety.ContentHarmEvaluator> provides single-shot evaluation for the four metrics supported by `HateAndUnfairnessEvaluator`, `SelfHarmEvaluator`, `ViolenceEvaluator`, and `SexualEvaluator`. ## Cached responses The library uses *response caching* functionality, which means responses from the AI model are persisted in a cache. In subsequent runs, if the request parameters (prompt and model) are unchanged, responses are then served from the cache to enable faster execution and lower cost. ## Reporting The library contains support for storing evaluation results and generating reports. The following image shows an example report in an Azure DevOps pipeline: The `dotnet aieval` tool, which ships as part of the `Microsoft.Extensions.AI.Evaluation.Console` package, includes functionality for generating reports and managing the stored evaluation data and cached responses. For more information, see [Generate a report](evaluate-with-reporting.md#generate-a-report). ## Configuration The libraries are designed to be flexible. You can pick the components that you need. For example, you can disable response caching or tailor reporting to work best in your environment. You can also customize and configure your evaluations, for example, by adding customized metrics and reporting options. ## Samples For a more comprehensive tour of the functionality and APIs available in the Microsoft.Extensions.AI.Evaluation libraries, see the [API usage examples (dotnet/ai-samples repo)](https://github.com/dotnet/ai-samples/blob/main/src/microsoft-extensions-ai-evaluation/api/). These examples are structured as a collection of unit tests. Each unit test showcases a specific concept or API and builds on the concepts and APIs showcased in previous unit tests. ## See also - [Evaluation of generative AI apps (Foundry)](/azure/ai-studio/concepts/evaluation-approach-gen-ai) -
responsible-ai.md 2.3 KB
--- title: Responsible AI with .NET description: Learn what responsible AI is and how you can use .NET to evaluate the safety of your AI apps. ms.date: 09/08/2025 ai-usage: ai-assisted --- # Responsible AI with .NET *Responsible AI* refers to the practice of designing, developing, and deploying artificial intelligence systems in a way that is ethical, transparent, and aligned with human values. It emphasizes fairness, accountability, privacy, and safety to ensure that AI technologies benefit individuals and society as a whole. As AI becomes increasingly integrated into applications and decision-making processes, prioritizing responsible AI is of utmost importance. Microsoft has identified [six principles](https://www.microsoft.com/ai/responsible-ai) for responsible AI: - Fairness - Reliability and safety - Privacy and security - Inclusiveness - Transparency - Accountability If you're building an AI app with .NET, the [📦 Microsoft.Extensions.AI.Evaluation.Safety](https://www.nuget.org/packages/Microsoft.Extensions.AI.Evaluation.Safety) package provides evaluators to help ensure that the responses your app generates, both text and image, meet the standards for responsible AI. The evaluators can also detect problematic content in user input. These safety evaluators use the [Microsoft Foundry Evaluation service](/azure/ai-foundry/concepts/evaluation-evaluators/risk-safety-evaluators) to perform evaluations. They include metrics for hate and unfairness, groundedness, ungrounded inference of human attributes, and the presence of: - Protected material - Self-harm content - Sexual content - Violent content - Vulnerable code (text-based only) - Indirect attacks (text-based only) For more information about the safety evaluators, see [Safety evaluators](libraries.md#safety-evaluators). To get started with the Microsoft.Extensions.AI.Evaluation.Safety evaluators, see [Tutorial: Evaluate response safety with caching and reporting](evaluate-safety.md). ## See also - [Responsible AI at Microsoft](https://www.microsoft.com/ai/responsible-ai) - [Training: Embrace responsible AI principles and practices](/training/modules/embrace-responsible-ai-principles-practices/) - [Foundry Evaluation service](/azure/ai-foundry/concepts/evaluation-evaluators/risk-safety-evaluators) - [Azure AI Content Safety](/azure/ai-services/content-safety/overview)
-
-
how-to
-
access-data-in-functions.md 5.7 KB
--- title: Access data in AI functions description: Learn how to pass data to AIFunction objects and how to access the data within the function delegate. ms.date: 11/17/2025 --- # Access data in AI functions When you create AI functions, you might need to access contextual data beyond the parameters provided by the AI model. The `Microsoft.Extensions.AI` library provides several mechanisms to pass data to function delegates. ## `AIFunction` class The <xref:Microsoft.Extensions.AI.AIFunction> type represents a function that can be described to an AI service and invoked. You can create `AIFunction` objects by calling one of the <xref:Microsoft.Extensions.AI.AIFunctionFactory.Create*?displayProperty=nameWithType> overloads. But <xref:Microsoft.Extensions.AI.AIFunction> is also a base class, and you can derive from it and implement your own AI function type. <xref:Microsoft.Extensions.AI.DelegatingAIFunction> provides an easy way to wrap an existing `AIFunction` and layer in additional functionality, including capturing additional data to be used. ## Pass data You can associate data with the function at the time it's created, either via closure or via <xref:Microsoft.Extensions.AI.ChatOptions.AdditionalProperties>. If you're creating your own function, you can populate `AdditionalProperties` however you want. If you use <xref:Microsoft.Extensions.AI.AIFunctionFactory> to create the function, you can populate data using <xref:Microsoft.Extensions.AI.AIFunctionFactoryOptions.AdditionalProperties?displayProperty=nameWithType>. You can also capture any references to data as part of the delegate provided to `AIFunctionFactory`. That is, you can bake in whatever you want to reference as part of the `AIFunction` itself. ## Access data in function delegates You might call your `AIFunction` directly, or you might call it indirectly by using <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient>. The following sections describe how to access argument data using either approach. ### Manual function invocation If you manually invoke an <xref:Microsoft.Extensions.AI.AIFunction> by calling <xref:Microsoft.Extensions.AI.AIFunction.InvokeAsync(Microsoft.Extensions.AI.AIFunctionArguments,System.Threading.CancellationToken)?displayProperty=nameWithType>, you pass in <xref:Microsoft.Extensions.AI.AIFunctionArguments>. The <xref:Microsoft.Extensions.AI.AIFunctionArguments> type includes: - A dictionary of named arguments. - <xref:Microsoft.Extensions.AI.AIFunctionArguments.Context>: An arbitrary `IDictionary<object, object>` for passing additional ambient data into the function. - <xref:Microsoft.Extensions.AI.AIFunctionArguments.Services>: An <xref:System.IServiceProvider> that lets the `AIFunction` resolve arbitrary state from a [dependency injection (DI)](../../core/extensions/dependency-injection/overview.md) container. If you want to access either the `AIFunctionArguments` or the `IServiceProvider` from within your <xref:Microsoft.Extensions.AI.AIFunctionFactory.Create*?displayProperty=nameWithType> delegate, create a parameter typed as `IServiceProvider` or `AIFunctionArguments`. That parameter will be bound to the relevant data from the `AIFunctionArguments` passed to `AIFunction.InvokeAsync()`. The following code shows an example: <xref:System.Threading.CancellationToken> is also special-cased: if the `AIFunctionFactory.Create` delegate or lambda has a `CancellationToken` parameter, it will be bound to the `CancellationToken` that was passed to `AIFunction.InvokeAsync()`. ### Invocation through `FunctionInvokingChatClient` <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient> publishes state about the current invocation to <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient.CurrentContext?displayProperty=nameWithType>, including not only the arguments, but all of the input `ChatMessage` objects, the <xref:Microsoft.Extensions.AI.ChatOptions>, and details on which function is being invoked (out of how many). You can add any data you want into <xref:Microsoft.Extensions.AI.ChatOptions.AdditionalProperties?displayProperty=nameWithType> and extract that inside of your `AIFunction` from `FunctionInvokingChatClient.CurrentContext.Options.AdditionalProperties`. The following code shows an example: #### Dependency injection If you use <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient> to invoke functions automatically, that client configures an <xref:Microsoft.Extensions.AI.AIFunctionArguments> object that it passes into the `AIFunction`. Because `AIFunctionArguments` includes the `IServiceProvider` that the `FunctionInvokingChatClient` was itself provided with, if you construct your client using standard DI means, that `IServiceProvider` is passed all the way into your `AIFunction`. At that point, you can query it for anything you want from DI. ## Advanced techniques If you want more fine-grained control over how parameters are bound, you can use <xref:Microsoft.Extensions.AI.AIFunctionFactoryOptions.ConfigureParameterBinding?displayProperty=nameWithType>, which puts you in control over how each parameter is populated. For example, the [MCP C# SDK uses this technique](https://github.com/modelcontextprotocol/csharp-sdk/blob/d344c651203841ec1c9e828736d234a6e4aebd07/src/ModelContextProtocol.Core/Server/AIFunctionMcpServerTool.cs#L83-L107) to automatically bind parameters from DI. If you use the <xref:Microsoft.Extensions.AI.AIFunctionFactory.Create(System.Reflection.MethodInfo,System.Func{Microsoft.Extensions.AI.AIFunctionArguments,System.Object},Microsoft.Extensions.AI.AIFunctionFactoryOptions)?displayProperty=nameWithType> overload, you can also run your own arbitrary logic when you create the target object that the instance method will be called on, each time. And you can do whatever you want to configure that instance. -
app-service-aoai-auth.md 7.9 KB
--- title: "Authenticate an Azure hosted .NET app to Azure OpenAI using Microsoft Entra ID" description: "Learn how to authenticate your Azure hosted .NET app to an Azure OpenAI resource using Microsoft Entra ID." author: alexwolfmsft ms.author: alexwolf ms.topic: how-to ms.date: 03/04/2026 ai-usage: ai-assisted zone_pivot_groups: azure-interface #customer intent: As a .NET developer, I want authenticate and authorize my App Service to Azure OpenAI by using Microsoft Entra so that I can securely use AI in my .NET application. --- # Authenticate to Azure OpenAI from an Azure hosted app using Microsoft Entra ID This article demonstrates how to use [Microsoft Entra ID managed identities](/azure/app-service/overview-managed-identity) and the [Microsoft.Extensions.AI library](../microsoft-extensions-ai.md) to authenticate an Azure hosted app to an Azure OpenAI resource. A managed identity from Microsoft Entra ID lets your app access other Microsoft Entra protected resources such as Azure OpenAI. Azure manages the identity and doesn't require you to provision, manage, or rotate any secrets. ## Prerequisites * An Azure account that has an active subscription. [Create an account for free](https://azure.microsoft.com/pricing/purchase-options/azure-account?cid=msft_learn). * [.NET SDK](https://dotnet.microsoft.com/download/visual-studio-sdks) * [Create and deploy an Azure OpenAI Service resource](/azure/ai-services/openai/how-to/create-resource) * [Create and deploy a .NET application to App Service](/azure/app-service/quickstart-dotnetcore) ## Add a managed identity to App Service Managed identities provide an automatically managed identity in Microsoft Entra ID for applications to use when connecting to resources that support Microsoft Entra authentication. Applications can use managed identities to obtain Microsoft Entra tokens without having to manage any credentials. You can assign two types of identities to your application: * A **system-assigned identity** is tied to your application and is deleted if your app is deleted. An app can have only one system-assigned identity. * A **user-assigned identity** is a standalone Azure resource that can be assigned to your app. An app can have multiple user-assigned identities. :::zone target="docs" pivot="azure-portal" # [System-assigned](#tab/system-assigned) 1. Go to your app's page in the [Azure portal](https://aka.ms/azureportal), and then scroll down to the **Settings** group. 1. Select **Identity**. 1. On the **System assigned** tab, toggle *Status* to **On**, and then select **Save**. > [!NOTE] > The preceding screenshot demonstrates this process on an Azure App Service, but the steps are similar on other hosts such as Azure Container Apps. ## [User-assigned](#tab/user-assigned) To add a user-assigned identity to your app, create the identity, and then add its resource identifier to your app config. 1. Create a user-assigned managed identity resource by following [these instructions](/azure/active-directory/managed-identities-azure-resources/how-to-manage-ua-identity-portal#create-a-user-assigned-managed-identity). 1. In the left navigation pane of your app's page, scroll down to the **Settings** group. 1. Select **Identity**. 1. Select **User assigned** > **Add**. 1. Locate the identity that you created earlier, select it, and then select **Add**. > [!IMPORTANT] > After you select **Add**, the app restarts. > [!NOTE] > The preceding screenshot demonstrates this process on an Azure App Service, but the steps are similar on other hosts such as Azure Container Apps. --- :::zone-end :::zone target="docs" pivot="azure-cli" ## [System-assigned](#tab/system-assigned) Run the `az webapp identity assign` command to create a system-assigned identity: ```azurecli az webapp identity assign --name <appName> --resource-group <groupName> ``` ## [User-assigned](#tab/user-assigned) 1. Create a user-assigned identity: ```azurecli az identity create --resource-group <groupName> --name <identityName> ``` 1. Assign the identity to your app: ```azurecli az webapp identity assign --resource-group <groupName> --name <appName> --identities <identityId> ``` --- :::zone-end ## Add an Azure OpenAI user role to the identity :::zone target="docs" pivot="azure-portal" 1. In the [Azure portal](https://aka.ms/azureportal), go to the scope that you want to grant **Azure OpenAI** access to. The scope can be a **Management group**, **Subscription**, **Resource group**, or a specific **Azure OpenAI** resource. 1. In the left navigation pane, select **Access control (IAM)**. 1. Select **Add**, then select **Add role assignment**. 1. On the **Role** tab, select the **Cognitive Services OpenAI User** role. 1. On the **Members** tab, select the managed identity. 1. On the **Review + assign** tab, select **Review + assign** to assign the role. :::zone-end :::zone target="docs" pivot="azure-cli" Use the Azure CLI to assign the Cognitive Services OpenAI User role to your managed identity at different scopes. # [Resource](#tab/resource) ```azurecli az role assignment create --assignee "<managedIdentityObjectID>" \ --role "Cognitive Services OpenAI User" \ --scope "/subscriptions/<subscriptionId>/resourcegroups/<resourceGroupName>/providers/<providerName>/<resourceType>/<resourceSubType>/<resourceName>" ``` # [Resource group](#tab/resource-group) ```azurecli az role assignment create --assignee "<managedIdentityObjectID>" \ --role "Cognitive Services OpenAI User" \ --scope "/subscriptions/<subscriptionId>/resourcegroups/<resourceGroupName>" ``` # [Subscription](#tab/subscription) ```azurecli az role assignment create --assignee "<managedIdentityObjectID>" \ --role "Cognitive Services OpenAI User" \ --scope "/subscriptions/<subscriptionId>" ``` # [Management group](#tab/management-group) ```azurecli az role assignment create --assignee "<managedIdentityObjectID>" \ --role "Cognitive Services OpenAI User" \ --scope "/providers/Microsoft.Management/managementGroups/<managementGroupName>" ``` --- :::zone-end ## Implement identity authentication in your app code 1. Add the following NuGet packages to your app: ```dotnetcli dotnet add package Azure.Identity dotnet add package Azure.AI.OpenAI dotnet add package Microsoft.Extensions.Azure dotnet add package Microsoft.Extensions.AI dotnet add package Microsoft.Extensions.AI.OpenAI ``` The preceding packages each handle the following concerns for this scenario: - **[Azure.Identity](https://www.nuget.org/packages/Azure.Identity)**: Provides core functionality to work with Microsoft Entra ID - **[Azure.AI.OpenAI](https://www.nuget.org/packages/Azure.AI.OpenAI)**: Lets your app interface with the Azure OpenAI service - **[Microsoft.Extensions.Azure](https://www.nuget.org/packages/Microsoft.Extensions.Azure)**: Provides helper extensions to register services for dependency injection - **[Microsoft.Extensions.AI](https://www.nuget.org/packages/Microsoft.Extensions.AI)**: Provides AI abstractions for common AI tasks - **[Microsoft.Extensions.AI.OpenAI](https://www.nuget.org/packages/Microsoft.Extensions.AI.OpenAI)**: Lets you use OpenAI service types as AI abstractions provided by **Microsoft.Extensions.AI** 1. In the `Program.cs` file of your app, create a `DefaultAzureCredential` object to discover and configure available credentials: 1. Create an AI service and register it with the service collection: 1. Inject the registered service for use in your endpoints: > [!TIP] > For more information about ASP.NET Core dependency injection and registering other AI service types, see the Azure SDK for .NET [dependency injection](../../azure/sdk/dependency-injection.md) documentation. ## Related content * [How to use managed identities for App Service and Azure Functions](/azure/app-service/overview-managed-identity) * [Role-based access control for Azure OpenAI Service](/azure/ai-services/openai/how-to/role-based-access-control) -
content-filtering.md 2.6 KB
--- title: "Manage Azure OpenAI content filtering in a .NET app" description: "Learn how to manage Azure OpenAI content filtering programmatically in a .NET app using the Azure OpenAI client library." ms.topic: how-to ms.date: 03/04/2026 ai-usage: ai-assisted --- # Work with Azure OpenAI content filtering in a .NET app This article shows how to handle content filtering in a .NET app. Azure OpenAI Service includes a content filtering system that works alongside core models. It runs both the prompt and completion through an ensemble of classification models to detect and take action on specific categories of potentially harmful content in both input prompts and output completions. Variations in API configurations and application design might affect completions and thus filtering behavior. For a deeper exploration of content filtering concepts and concerns, see the [Content Filtering](/azure/ai-services/openai/concepts/content-filter) documentation. ## Prerequisites * An Azure account that has an active subscription. [Create an account for free](https://azure.microsoft.com/pricing/purchase-options/azure-account?cid=msft_learn). * [.NET SDK](https://dotnet.microsoft.com/download/visual-studio-sdks) * [Create and deploy an Azure OpenAI Service resource](/azure/ai-services/openai/how-to/create-resource) ## Configure and test the content filter To use the sample code in this article, you need to create and assign a content filter to your OpenAI model. 1. [Create and assign a content filter](/azure/ai-services/openai/how-to/content-filters) to your provisioned model. 1. Add the [`Azure.AI.OpenAI`](https://www.nuget.org/packages/Azure.AI.OpenAI) NuGet package to your project. ```dotnetcli dotnet add package Azure.AI.OpenAI ``` Or, in .NET 10+: ```dotnetcli dotnet package add Azure.AI.OpenAI ``` 1. Create a simple chat completion flow in your .NET app using the `AzureOpenAiClient`. Replace the `YOUR_MODEL_ENDPOINT` and `YOUR_MODEL_DEPLOYMENT_NAME` values with your own. 1. Replace the `YOUR_PROMPT` placeholder with your own message and run the app to experiment with content filtering results. If you enter a prompt the AI considers unsafe, Azure OpenAI returns a `400 Bad Request` code. The app prints a message in the console similar to the following: ```output The response was filtered due to the prompt triggering Azure OpenAI's content management policy... ``` ## Related content * [Create and assign a content filter](/azure/ai-services/openai/how-to/content-filters) * [Content Filtering concepts](/azure/ai-services/openai/concepts/content-filter) * [Create a chat app](../quickstarts/prompt-model.md) -
handle-invalid-tool-input.md 3.9 KB
--- title: Handle invalid function input from AI models description: Learn strategies to handle invalid function input when AI models provide incorrect or malformed function call parameters. ms.date: 01/05/2026 ai-usage: ai-assisted --- # Handle invalid function input from AI models When AI models call functions in your .NET code, they might sometimes provide invalid input that doesn't match the expected schema. The `Microsoft.Extensions.AI` library provides several strategies to handle these scenarios gracefully. ## Common scenarios for invalid input AI models can provide invalid function call input in several ways: - Missing required parameters. - Incorrect data types (for example, sending a string when an integer is expected). - Malformed JSON that can't be deserialized. Without proper error handling, these issues can cause your application to fail or provide poor user experiences. ## Enable detailed error messages By default, when a function invocation fails, the AI model receives a generic error message. You can enable detailed error reporting using the <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient.IncludeDetailedErrors?displayProperty=nameWithType> property. When this property is set to `true` and an error occurs during function invocation, the full exception message is added to the chat history. This allows the AI model to see what went wrong and potentially self-correct in subsequent attempts. > [!NOTE] > Setting `IncludeDetailedErrors` to `true` can expose internal system details to the AI model and potentially to end users. Ensure exception messages don't contain secrets, connection strings, or other sensitive information. To avoid leaking sensitive information, consider disabling detailed errors in production environments. ## Implement custom error handling For more control over error handling, you can set a custom <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient.FunctionInvoker?displayProperty=nameWithType> delegate. This allows you to intercept function calls, catch exceptions, and return custom error messages to the AI model. The following example shows how to implement a custom function invoker that catches serialization errors and provides helpful feedback: By returning descriptive error messages instead of throwing exceptions, you allow the AI model to see what went wrong and try again with corrected input. ### Best practices for error messages When returning error messages to enable AI self-correction, provide clear, actionable feedback: - **Be specific**: Explain exactly what was wrong with the input. - **Provide examples**: Show the expected format or valid values. - **Use consistent format**: Help the AI model learn from patterns. - **Log errors**: Track error patterns for debugging and monitoring. ## Use strict JSON schema (OpenAI only) When using OpenAI models, you can enable strict JSON schema mode to enforce that the model's output strictly adheres to your function's schema. This helps prevent type mismatches and missing required fields. Enable strict mode using the `Strict` additional property on your function metadata. When enabled, OpenAI models try to ensure their output matches your schema exactly: For the latest list of models that support strict JSON schema, check the [OpenAI documentation](https://platform.openai.com/docs/guides/structured-outputs). ### Limitations While strict mode significantly improves schema adherence, keep these limitations in mind: - Not all JSON Schema features are supported in strict mode. - Complex schemas might still produce occasional errors. - Always validate outputs even with strict mode enabled. - Strict mode is OpenAI-specific and doesn't apply to other AI providers. ## Next steps - [Access data in AI functions](access-data-in-functions.md) - [Execute a local .NET function](../quickstarts/use-function-calling.md) - [Build a chat app](../quickstarts/build-chat-app.md) -
use-tokenizers.md 4.9 KB
--- title: Use Microsoft.ML.Tokenizers for text tokenization description: Learn how to use the Microsoft.ML.Tokenizers library to tokenize text for AI models, manage token counts, and work with various tokenization algorithms. ms.topic: how-to ms.date: 10/29/2025 ai-usage: ai-assisted --- # Use Microsoft.ML.Tokenizers for text tokenization The [Microsoft.ML.Tokenizers](https://www.nuget.org/packages/Microsoft.ML.Tokenizers) library provides a comprehensive set of tools for tokenizing text in .NET applications. Tokenization is essential when you work with large language models (LLMs), as it allows you to manage token counts, estimate costs, and preprocess text for AI models. This article shows you how to use the library's key features and work with different tokenizer models. ## Prerequisites - [.NET 8 SDK](https://dotnet.microsoft.com/download/dotnet/8.0) or later > [!NOTE] > The Microsoft.ML.Tokenizers library also supports .NET Standard 2.0, making it compatible with .NET Framework 4.6.1 and later. ## Install the package Install the Microsoft.ML.Tokenizers NuGet package: ```dotnetcli dotnet add package Microsoft.ML.Tokenizers ``` For Tiktoken models (like GPT-4), you also need to install the corresponding data package: ```dotnetcli dotnet add package Microsoft.ML.Tokenizers.Data.O200kBase ``` ## Key features The Microsoft.ML.Tokenizers library provides: - **Extensible tokenizer architecture**: Allows specialization of Normalizer, PreTokenizer, Model/Encoder, and Decoder components. - **Multiple tokenization algorithms**: Supports BPE (byte-pair encoding), Tiktoken, Llama, CodeGen, and more. - **Token counting and estimation**: Helps manage costs and context limits when working with AI services. - **Flexible encoding options**: Provides methods to encode text to token IDs, count tokens, and decode tokens back to text. ## Use Tiktoken tokenizer The Tiktoken tokenizer is commonly used with OpenAI models like GPT-4. The following example shows how to initialize a Tiktoken tokenizer and perform common operations: For better performance, you should cache and reuse the tokenizer instance throughout your app. When you work with LLMs, you often need to manage text within token limits. The following example shows how to trim text to a specific token count: ## Use Llama tokenizer The Llama tokenizer is designed for the Llama family of models. It requires a tokenizer model file, which you can download from model repositories like Hugging Face: All tokenizers support advanced encoding options, such as controlling normalization and pretokenization: ## Use BPE tokenizer *Byte-pair encoding* (BPE) is the underlying algorithm used by many tokenizers, including Tiktoken. BPE was initially developed as an algorithm to compress texts, and then used by OpenAI for tokenization when it pretrained the GPT model. The following example demonstrates BPE tokenization: The library also provides specialized tokenizers like <xref:Microsoft.ML.Tokenizers.BpeTokenizer> and <xref:Microsoft.ML.Tokenizers.EnglishRobertaTokenizer> that you can configure with custom vocabularies for specific models. For more information about BPE, see [Byte-pair encoding tokenization](https://huggingface.co/learn/llm-course/chapter6/5). ## Common tokenizer operations All tokenizers in the library implement the <xref:Microsoft.ML.Tokenizers.Tokenizer> base class. The following table shows the available methods. | Method | Description | |-------------------------------------------------------|--------------------------------------| | <xref:Microsoft.ML.Tokenizers.Tokenizer.EncodeToIds*> | Converts text to a list of token IDs. | | <xref:Microsoft.ML.Tokenizers.Tokenizer.Decode*> | Converts token IDs back to text. | | <xref:Microsoft.ML.Tokenizers.Tokenizer.CountTokens*> | Returns the number of tokens in a text string. | | <xref:Microsoft.ML.Tokenizers.Tokenizer.EncodeToTokens*> | Returns detailed token information including values and IDs. | | <xref:Microsoft.ML.Tokenizers.Tokenizer.GetIndexByTokenCount*> | Finds the character index for a specific token count from the start. | | <xref:Microsoft.ML.Tokenizers.Tokenizer.GetIndexByTokenCountFromEnd*> | Finds the character index for a specific token count from the end. | ## Migrate from other libraries If you're currently using `DeepDev.TokenizerLib` or `SharpToken`, consider migrating to Microsoft.ML.Tokenizers. The library has been enhanced to cover scenarios from those libraries and provides better performance and support. For migration guidance, see the [migration guide](https://github.com/dotnet/machinelearning/blob/main/docs/code/microsoft-ml-tokenizers-migration-guide.md). ## Related content - [Understanding tokens](../conceptual/understanding-tokens.md) - [Microsoft.ML.Tokenizers API reference](/dotnet/api/microsoft.ml.tokenizers) - [Microsoft.ML.Tokenizers NuGet package](https://www.nuget.org/packages/Microsoft.ML.Tokenizers)
-
-
quickstarts
-
ai-templates.md 1.8 KB
--- title: Quickstart - Create a .NET AI app using the AI app template description: Create a .NET AI app to chat with custom data using the AI app template extensions and the Microsoft.Extensions.AI libraries ms.date: 03/04/2026 ms.topic: quickstart zone_pivot_groups: meai-targets ai-usage: ai-assisted --- # Create a .NET AI app to chat with custom data using the AI app template extensions In this quickstart, you learn how to create a .NET AI app to chat with custom data using the .NET AI app template. The template is designed to streamline the getting started experience for building AI apps with .NET by handling common setup tasks and configurations for you. :::zone target="docs" pivot="github-models" [!INCLUDE [ai-templates-github-models](includes/ai-templates-github-models.md)] :::zone-end :::zone target="docs" pivot="azure-openai" [!INCLUDE [ai-templates-azure-openai](includes/ai-templates-azure-openai.md)] :::zone-end :::zone target="docs" pivot="openai" [!INCLUDE [ai-templates-openai](includes/ai-templates-openai.md)] :::zone-end :::zone target="docs" pivot="ollama" [!INCLUDE [ai-templates-ollama](includes/ai-templates-ollama.md)] :::zone-end ## Run and test the app 1. Select the run button at the top of Visual Studio to launch the app. After a moment, you should see the following UI load in the browser: 1. Enter a prompt into the input box such as *"What are some essential tools in the survival kit?"* to ask your AI model a question about the ingested data from the example files. The app responds with an answer to the question and provides citations of where it found the data. You can click on one of the citations to be directed to the relevant section of the example files. ## Next steps - [Generate text and conversations with .NET and Azure OpenAI Completions](/training/modules/open-ai-dotnet-text-completions/) -
build-chat-app.md 4.4 KB
--- title: Quickstart - Build an AI chat app with .NET description: Create a simple AI powered chat app using Microsoft.Extensions.AI and the OpenAI or Azure OpenAI SDKs ms.date: 03/04/2026 ms.topic: quickstart zone_pivot_groups: openai-library ai-usage: ai-assisted --- # Build an AI chat app with .NET In this quickstart, you learn how to create a conversational .NET console chat app using an OpenAI or Azure OpenAI model. The app uses the <xref:Microsoft.Extensions.AI> library so you can write code using AI abstractions rather than a specific SDK. AI abstractions enable you to change the underlying AI model with minimal code changes. :::zone target="docs" pivot="openai" [!INCLUDE [openai-prereqs](includes/prerequisites-openai.md)] :::zone-end :::zone target="docs" pivot="azure-openai" [!INCLUDE [azure-openai-prereqs](includes/prerequisites-azure-openai.md)] :::zone-end ## Create the app Complete the following steps to create a .NET console app to connect to an AI model. 1. In an empty directory on your computer, use the `dotnet new` command to create a new console app: ```dotnetcli dotnet new console -o ChatAppAI ``` 1. Change directory into the app folder: ```dotnetcli cd ChatAppAI ``` 1. Install the required packages: :::zone target="docs" pivot="azure-openai" ```bash dotnet add package Azure.Identity dotnet add package Azure.AI.OpenAI dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` :::zone-end :::zone target="docs" pivot="openai" ```bash dotnet add package OpenAI dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` :::zone-end 1. Open the app in Visual Studio Code (or your editor of choice). ```bash code . ``` :::zone target="docs" pivot="azure-openai" [!INCLUDE [create-ai-service](includes/create-ai-service.md)] :::zone-end :::zone target="docs" pivot="openai" ## Configure the app 1. Navigate to the root of your .NET project from a terminal or command prompt. 1. Run the following commands to configure your OpenAI API key as a secret for the sample app: ```bash dotnet user-secrets init dotnet user-secrets set OpenAIKey <your-OpenAI-key> dotnet user-secrets set ModelName <your-OpenAI-model-name> ``` :::zone-end ## Add the app code This app uses the [`Microsoft.Extensions.AI`](https://www.nuget.org/packages/Microsoft.Extensions.AI/) package to send and receive requests to the AI model. The app provides users with information about hiking trails. 1. In the `Program.cs` file, add the following code to connect and authenticate to the AI model. :::zone target="docs" pivot="azure-openai" > [!NOTE] > <xref:Azure.Identity.DefaultAzureCredential> searches for authentication credentials from your local tooling. If you aren't using the `azd` template to provision the Azure OpenAI resource, you'll need to assign the `Azure AI Developer` role to the account you used to sign in to Visual Studio or the Azure CLI. For more information, see [Authenticate to Foundry tools with .NET](../azure-ai-services-authentication.md). :::zone-end :::zone target="docs" pivot="openai" :::zone-end 1. Create a system prompt to provide the AI model with initial role context and instructions about hiking recommendations: 1. Create a conversational loop that accepts an input prompt from the user, sends the prompt to the model, and prints the response completion: 1. Use the `dotnet run` command to run the app: ```dotnetcli dotnet run ``` The app prints out the completion response from the AI model. Send additional follow up prompts and ask other questions to experiment with the AI chat functionality. :::zone target="docs" pivot="azure-openai" ## Clean up resources If you no longer need them, delete the Azure OpenAI resource and GPT-4 model deployment. 1. In the [Azure portal](https://aka.ms/azureportal), navigate to the Azure OpenAI resource. 1. Select the Azure OpenAI resource, and then select **Delete**. :::zone-end ## Next steps - [Quickstart - Chat with a local AI model](chat-local-model.md) - [Generate images from text using AI](text-to-image.md) -
build-mcp-client.md 3 KB
--- title: Quickstart - Create a minimal MCP client using .NET description: Learn to create a minimal MCP client and connect it to an MCP server using .NET ms.date: 03/04/2026 ms.topic: quickstart author: alexwolfmsft ai-usage: ai-assisted --- # Create a minimal MCP client using .NET In this quickstart, you build a minimal [Model Context Protocol (MCP)](../get-started-mcp.md) client using the [C# SDK for MCP](https://github.com/modelcontextprotocol/csharp-sdk). You also learn how to configure the client to connect to an MCP server, such as the one created in the [Build a minimal MCP server](build-mcp-server.md) quickstart. ## Prerequisites - [.NET 8.0 SDK or higher](https://dotnet.microsoft.com/download) - [Visual Studio Code](https://code.visualstudio.com/) > [!NOTE] > The MCP client you build in the sections ahead connects to the sample MCP server from the [Build a minimal MCP server](build-mcp-server.md) quickstart. You can also use your own MCP server if you provide your own connection configuration. ## Create the .NET host app Complete the following steps to create a .NET console app. The app acts as a host for an MCP client that connects to an MCP server. ### Create the project 1. In a terminal window, navigate to the directory where you want to create your app, and create a new console app with the `dotnet new` command: ```console dotnet new console -n MCPHostApp ``` 1. Navigate into the newly created project folder: ```console cd MCPHostApp ``` 1. Run the following commands to add the necessary NuGet packages: ```console dotnet add package Azure.AI.OpenAI --prerelease dotnet add package Azure.Identity dotnet add package Microsoft.Extensions.AI dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package ModelContextProtocol --prerelease ``` 1. Open the project folder in your editor of choice, such as Visual Studio Code: ```console code . ``` ### Add the app code Replace the contents of `Program.cs` with the following code: The preceding code accomplishes the following tasks: - Initializes an `IChatClient` abstraction using the [`Microsoft.Extensions.AI`](../microsoft-extensions-ai.md) libraries. - Creates an MCP client and configures it to connect to your MCP server. - Retrieves and displays a list of available tools from the MCP server, which is a standard MCP function. - Implements a conversational loop that processes user prompts and utilizes the tools for responses. ## Run and test the app Complete the following steps to test your .NET host app: 1. In a terminal window open to the root of your project, run the following command to start the app: ```console dotnet run ``` 1. Once the app is running, enter a prompt to run the **ReverseEcho** tool: ```console Reverse the following: "Hello, minimal MCP server!" ``` 1. Verify that the server responds with the echoed message: ```output !revres PCM laminim ,olleH ``` ## Related content [Get started with .NET AI and the Model Context Protocol](../get-started-mcp.md) -
build-mcp-server.md 21.3 KB
--- title: Quickstart - Create a minimal MCP server and publish to NuGet description: Learn to create and connect to a minimal MCP server using C# and publish it to NuGet. ms.date: 03/04/2026 ms.topic: quickstart author: alexwolfmsft ms.author: alexwolf zone_pivot_groups: development-environment-one ai-usage: ai-assisted --- # Create a minimal MCP server using C# and publish to NuGet In this quickstart, you create a minimal Model Context Protocol (MCP) server using the [C# SDK for MCP](https://github.com/modelcontextprotocol/csharp-sdk), connect to it using GitHub Copilot, and publish it to NuGet (stdio transport only). MCP servers are services that expose capabilities to clients through the Model Context Protocol (MCP). > [!NOTE] > The `Microsoft.McpServer.ProjectTemplates` template package is currently in preview. ## Prerequisites ::: zone pivot="visualstudio" - [.NET 10.0 SDK](https://dotnet.microsoft.com/download/dotnet) - [Visual Studio 2022 or higher](https://visualstudio.microsoft.com/) - [GitHub Copilot](https://github.com/features/copilot) - [NuGet.org account](https://www.nuget.org/users/account/LogOn) ::: zone-end ::: zone pivot="vscode" - [.NET 10.0 SDK](https://dotnet.microsoft.com/download/dotnet) - [Visual Studio Code](https://code.visualstudio.com/) (optional) - [C# Dev Kit extension](https://marketplace.visualstudio.com/items?itemName=ms-dotnettools.csdevkit) - [GitHub Copilot extension](https://marketplace.visualstudio.com/items?itemName=GitHub.copilot) for Visual Studio Code - [NuGet.org account](https://www.nuget.org/users/account/LogOn) ::: zone-end ::: zone pivot="cli" - [.NET 10.0 SDK](https://dotnet.microsoft.com/download/dotnet) - [Visual Studio Code](https://code.visualstudio.com/) (optional) - [Visual Studio](https://visualstudio.microsoft.com/) (optional) - [GitHub Copilot](https://github.com/features/copilot) / [GitHub Copilot extension](https://marketplace.visualstudio.com/items?itemName=GitHub.copilot) for Visual Studio Code - [NuGet.org account](https://www.nuget.org/users/account/LogOn) ::: zone-end ## Create the project ::: zone pivot="visualstudio" 1. In a terminal window, install the MCP Server template: ```dotnetcli dotnet new install Microsoft.McpServer.ProjectTemplates ``` > [!NOTE] > .NET 10.0 SDK or a later version is required to install `Microsoft.McpServer.ProjectTemplates`. 1. Open Visual Studio, and select **Create a new project** in the start window (or select **File** > **New** > **Project/Solution** from inside Visual Studio). 1. In the **Create a new project** window, select **C#** from the Language list and **AI** from the **All project types** list. After you apply the language and project type filters, select the **MCP Server App** template, and then select **Next**. 1. In the **Configure your new project** window, enter **MyMcpServer** in the **Project name** field. Then, select **Next**. 1. In the **Additional information** window, you can configure the following options: - **Framework**: Select the target .NET framework. - **MCP Server Transport Type**: Choose between creating a **local** (stdio) or a **remote** (http) MCP server. - **Enable native AOT (Ahead-Of-Time) publish**: Enable your MCP server to be self-contained and compiled to native code. For more information, see the [Native AOT deployment guide](../../core/deploying/native-aot/index.md). - **Enable self-contained publish**: Enable your MCP server to be published as a self-contained executable. For more information, see the [Self-contained deployment section of the .NET application publishing guide](../../core/deploying/index.md#self-contained-deployment). Choose your preferred options or keep the default ones, and then select **Create**. Visual Studio opens your new project. 1. Update the `<PackageId>` in the `.csproj` file to be unique on NuGet.org, for example `<NuGet.org username>.SampleMcpServer`. ::: zone-end ::: zone pivot="vscode" 1. In a terminal window, install the MCP Server template: ```bash dotnet new install Microsoft.McpServer.ProjectTemplates ``` > [!NOTE] > .NET 10.0 SDK or later is required to install `Microsoft.McpServer.ProjectTemplates`. 1. Open Visual Studio Code. 1. Go to the **Explorer** view and select **Create .NET Project**. Alternatively, you can bring up the Command Palette using <kbd>Ctrl+Shift+P</kbd> (<kbd>Command+Shift+P</kbd> on MacOS) and then type ".NET" to find and select the **.NET: New Project** command. This action will bring up a dropdown list of .NET projects. 1. After selecting the command, use the Search bar in the Command Palette or scroll down to locate the **MCP Server App** template. 1. Select the location where you would like the new project to be created. 1. Give your new project a name, **MyMCPServer**. Press **Enter**. 1. Select your solution file format (`.sln` or `.slnx`). 1. Select **Template Options**. Here, you can configure the following options: - **Framework**: Select the target .NET framework. - **MCP Server Transport Type**: Choose between creating a **local** (stdio) or a **remote** (http) MCP server. - **Enable native AOT (Ahead-Of-Time) publish**: Enable your MCP server to be self-contained and compiled to native code. For more information, see the [Native AOT deployment guide](../../core/deploying/native-aot/index.md). - **Enable self-contained publish**: Enable your MCP server to be published as a self-contained executable. For more information, see the [Self-contained deployment section of the .NET application publishing guide](../../core/deploying/index.md#self-contained-deployment). Choose your preferred options or keep the default ones, and then select **Create Project**. VS Code opens your new project. 1. Update the `<PackageId>` in the `.csproj` file to be unique on NuGet.org, for example `<NuGet.org username>.SampleMcpServer`. ::: zone-end ::: zone pivot="cli" 1. Create a new MCP server app with the `dotnet new mcpserver` command: ```bash dotnet new mcpserver -n SampleMcpServer ``` By default, this command creates a self-contained tool package targeting all of the most common platforms that .NET is supported on. To see more options, use `dotnet new mcpserver --help`. Using the `dotnet new mcpserver --help` command gives you several template options you can add when creating a new MCP server: - **Framework**: Select the target .NET framework. - **MCP Server Transport Type**: Choose between creating a **local** (stdio) or a **remote** (http) MCP server. - **Enable native AOT (Ahead-Of-Time) publish**: Enable your MCP server to be self-contained and compiled to native code. For more information, see the [Native AOT deployment guide](../../core/deploying/native-aot/index.md). - **Enable self-contained publish**: Enable your MCP server to be published as a self-contained executable. For more information, see the [Self-contained deployment section of the .NET application publishing guide](../../core/deploying/index.md#self-contained-deployment). 1. Navigate to the `SampleMcpServer` directory: ```bash cd SampleMcpServer ``` 1. Build the project: ```bash dotnet build ``` 1. Update the `<PackageId>` in the `.csproj` file to be unique on NuGet.org, for example `<NuGet.org username>.SampleMcpServer`. ::: zone-end ## Tour the MCP Server Project Creating your MCP server project via the template gives you the following major files: * `Program.cs`: A file defining the application as an MCP server and registering MCP services such as transport type and MCP tools. * Choosing the (default) **stdio** transport option in when creating the project, this file will be configured to define the MCP Server as a local one (that is, `.withStdioServerTransport()`). * Choosing the **http** transport option will configure this file to include remote transport-specific definitions (that is, `.withHttpServerTransport()`, `MapMcp()`). * `RandomNumberTools.cs`: A class defining an example MCP server tool that returns a random number between user-specified min/max values. * **[HTTP Transport Only]** `[MCPServerName].http`: A file defining the default host address for an HTTP MCP server and JSON-RPC communication. * `server.json`: A file defining how and where your MCP server is published. ::: zone pivot="visualstudio" ::: zone-end ::: zone pivot="cli,vscode" ::: zone-end ## Configure the MCP server ::: zone pivot="visualstudio" Configure GitHub Copilot for Visual Studio to use your custom MCP server. 1. In Visual Studio, select the GitHub Copilot icon in the top right corner and select **Open Chat Window**. 1. In the GitHub Copilot Chat window, click the **Select Tools** wrench icon followed by the plus icon in the top right corner. 1. In the **Add Custom MCP Server** dialog window, enter the following info: * **Destination**: Choose the scope of where your MCP server is configured: * **Solution** - The MCP server is available only across the active solution. * **Global** - The MCP server is available across all solutions. * **Server ID**: The unique name / identifier for your MCP server. * **Type**: The transport type of your MCP server (stdio or HTTP). * **Command (Stdio transport only)**: The command to run your stdio MCP server (that is, `dotnet run --project [relative path to .csproj file]`) * **URL (HTTP transport only)**: The address of your HTTP MCP server * **Environment Variables (optional)** 1. Select **Save**. A `.mcp.json` file will be added to the specified destination. **Stdio Transport `.mcp.json`** Add the relative path to your `.csproj` file under the "args" field. ```json { "inputs": [], "servers": { "MyMcpServer": { "type": "stdio", "command": "dotnet", "args": [ "run", "--project", "<relative-path-to-project-file>" ] } } } ``` **HTTP Transport `.mcp.json`** ```json { "inputs": [], "servers": { "MyMCPServer": { "url": "http://localhost:6278", "type": "http", "headers": {} } } } ``` ::: zone-end ::: zone pivot="vscode,cli" Configure GitHub Copilot for Visual Studio Code to use your custom MCP server, either via the VS Code Command Palette or manually. ### Command Palette configuration 1. Open the Command Palette using <kbd>Ctrl+Shift+P</kbd> (<kbd>Command+Shift+P</kbd> on macOS). Search "mcp" to locate the `MCP: Add Server` command. 1. Select the type of MCP server to add (typically the transport type you selected at project creation). 1. If adding a **stdio** MCP server, enter a command and optional arguments. For this example, use `dotnet run --project`. If adding an **HTTP** MCP server, enter the localhost or web address. 1. Enter a unique server ID (example: "MyMCPServer"). 1. Select a configuration target: * **Global**: Make the MCP server available across all workspaces. The generated `mcp.json` file will appear under your global user configuration. * **Workspace**: Make the MCP server available only from within the current workspace. The generated `mcp.json` file will appear under the `.vscode` folder within your workspace. 1. After you complete the previous steps, an `.mcp.json` file will be created in the location specified by the configuration target. **Stdio Transport `mcp.json`** Add the relative path to your `.csproj` file under the "args" field. ```json { "servers": { "MyMcpServer": { "type": "stdio", "command": "dotnet", "args": [ "run", "--project", "<relative-path-to-project-file>" ] } } } ``` **HTTP Transport `mcp.json`** ```json { "servers": { "MyMCPServer": { "url": "http://localhost:6278", "type": "http" } }, "inputs": [] } ``` ### Manual configuration 1. Create a `.vscode` folder at the root of your project. 1. Add an `mcp.json` file in the `.vscode` folder with the following content: ```json { "servers": { "SampleMcpServer": { "type": "stdio", "command": "dotnet", "args": [ "run", "--project", "<relative-path-to-project-file>" ] } } } ``` > [!NOTE] > VS Code executes MCP servers from the workspace root. The `<relative-path-to-project-file>` placeholder should point to your .NET project file. For example, the value for this **SampleMcpServer** app would be `SampleMcpServer.csproj`. 1. Save the file. ::: zone-end ## Test the MCP server ::: zone pivot="visualstudio" The MCP server template includes a tool called `get_random_number` you can use for testing and as a starting point for development. 1. Open GitHub Copilot chat in Visual Studio and switch to **Agent** mode. 1. Select the **Select tools** icon to verify your **MyMCPServer** is available with the sample tool listed. 1. Enter a prompt to run the **get_random_number** tool: ```console Give me a random number between 1 and 100. ``` 1. GitHub Copilot requests permission to run the **get_random_number** tool for your prompt. Select **Continue** or use the arrow to select a more specific behavior: - **Current session** always runs the operation in the current GitHub Copilot Agent Mode session. - **Current solution** always runs the command for the current VS solution. - **Always allow** sets the operation to always run for any GitHub Copilot Agent Mode session. 1. Verify that the server responds with a random number: ```output Your random number is 42. ``` ::: zone-end ::: zone pivot="vscode,cli" The MCP server template includes a tool called `get_random_number` you can use for testing and as a starting point for development. 1. Open GitHub Copilot chat in VS Code and switch to **Agent** mode. 1. Select the **Select tools** icon to verify your **MyMCPServer** is available with the sample tool listed. 1. Enter a prompt to run the **get_random_number** tool: ```console Give me a random number between 1 and 100. ``` 1. GitHub Copilot requests permission to run the **get_random_number** tool for your prompt. Select **Continue** or use the arrow to select a more specific behavior: - **Current session** always runs the operation in the current GitHub Copilot Agent Mode session. - **Current workspace** always runs the command for the current VS Code workspace. - **Always allow** sets the operation to always run for any GitHub Copilot Agent Mode session. 1. Verify that the server responds with a random number: ```output Your random number is 42. ``` ::: zone-end ## Add inputs and configuration options In this example, you enhance the MCP server to use a configuration value set in an environment variable. This could be configuration needed for the functioning of your MCP server, such as an API key, an endpoint to connect to, or a local directory path. 1. Add another tool method after the `GetRandomNumber` method in `Tools/RandomNumberTools.cs`. Update the tool code to use an environment variable. 1. Update the `.vscode/mcp.json` to set the `WEATHER_CHOICES` environment variable for testing. ```json { "servers": { "SampleMcpServer": { "type": "stdio", "command": "dotnet", "args": [ "run", "--project", "<relative-path-to-project-file>" ], "env": { "WEATHER_CHOICES": "sunny,humid,freezing" } } } } ``` 1. Try another prompt with Copilot in VS Code, such as: ```console What is the weather in Redmond, Washington? ``` VS Code should return a random weather description. 1. Update the `.mcp/server.json` to declare your environment variable input. The `server.json` file schema is defined by the [MCP Registry project](https://github.com/modelcontextprotocol/registry/blob/main/docs/reference/server-json/generic-server-json.md) and is used by NuGet.org to generate VS Code MCP configuration. - Use the `environmentVariables` property to declare environment variables used by your app that will be set by the client using the MCP server (for example, VS Code). - Use the `packageArguments` property to define CLI arguments that will be passed to your app. For more examples, see the [MCP Registry project](https://github.com/modelcontextprotocol/registry/blob/main/docs/reference/server-json/generic-server-json.md#examples). The only information used by NuGet.org in the `server.json` is the first `packages` array item with the `registryType` value matching `nuget`. The other top-level properties aside from the `packages` property are currently unused and are intended for the upcoming central MCP Registry. You can leave the placeholder values until the MCP Registry is live and ready to accept MCP server entries. You can [test your MCP server again](#test-the-mcp-server) before moving forward. ## Pack and publish to NuGet 1. Pack the project: ```bash dotnet pack -c Release ``` This command produces one tool package and several platform-specific packages based on the `<RuntimeIdentifiers>` list in `SampleMcpServer.csproj`. 1. Publish the packages to NuGet: ```bash dotnet nuget push bin/Release/*.nupkg --api-key <your-api-key> --source https://api.nuget.org/v3/index.json ``` Be sure to publish all `.nupkg` files to ensure every supported platform can run the MCP server. If you want to test the publishing flow before publishing to NuGet.org, you can register an account on the NuGet Gallery integration environment: [https://int.nugettest.org](https://int.nugettest.org). The `push` command would be modified to: ```bash dotnet nuget push bin/Release/*.nupkg --api-key <your-api-key> --source https://apiint.nugettest.org/v3/index.json ``` For more information, see [Publish a package](/nuget/nuget-org/publish-a-package). ## Discover MCP servers on NuGet.org 1. Search for your MCP server package on [NuGet.org](https://www.nuget.org/packages?packagetype=mcpserver) (or [int.nugettest.org](https://int.nugettest.org/packages?packagetype=mcpserver) if you published to the integration environment) and select it from the list. 1. View the package details and copy the JSON from the "MCP Server" tab. 1. In your `mcp.json` file in the `.vscode` folder, add the copied JSON, which looks like this: ```json { "inputs": [ { "type": "promptString", "id": "weather_choices", "description": "Comma separated list of weather descriptions to randomly select.", "password": false } ], "servers": { "Contoso.SampleMcpServer": { "type": "stdio", "command": "dnx", "args": ["Contoso.SampleMcpServer@0.0.1-beta", "--yes"], "env": { "WEATHER_CHOICES": "${input:weather_choices}" } } } } ``` If you published to the NuGet Gallery integration environment, you need to add `"--add-source", "https://apiint.nugettest.org/v3/index.json"` at the end of the `"args"` array. 1. Save the file. 1. In GitHub Copilot, select the **Select tools** icon to verify your **SampleMcpServer** is available with the tools listed. 1. Enter a prompt to run the new **get_city_weather** tool: ```console What is the weather in Redmond? ``` 1. If you added inputs to your MCP server (for example, `WEATHER_CHOICES`), you will be prompted to provide values. 1. Verify that the server responds with the random weather: ```output The weather in Redmond is balmy. ``` ## Common issues ### The command "dnx" needed to run SampleMcpServer was not found If VS Code shows this error when starting the MCP server, you need to install a compatible version of the .NET SDK. The `dnx` command is shipped as part of the .NET SDK, starting with version 10. [Install the .NET 10 SDK](https://dotnet.microsoft.com/download/dotnet) to resolve this issue. ### GitHub Copilot doesn't use your tool (an answer is provided without invoking your tool) Generally speaking, an AI agent like GitHub Copilot is informed that it has some tools available by the client application, such as VS Code. Some tools, such as the sample random number tool, might not be leveraged by the AI agent because it has similar functionality built in. If your tool is not being used, check the following: 1. Verify that your tool appears in the list of tools that VS Code has enabled. See the screenshot in [Test the MCP server](#test-the-mcp-server) for how to check this. 1. Explicitly reference the name of the tool in your prompt. In VS Code, you can reference your tool by name. For example, `Using #get_random_weather, what is the weather in Redmond?`. 1. Verify your MCP server is able to start. You can check this by clicking the "Start" button visible above your MCP server configuration in the VS Code user or workspace settings. ## Related content - [Get started with .NET AI and the Model Context Protocol](../get-started-mcp.md) - [Model Context Protocol .NET samples](https://github.com/microsoft/mcp-dotnet-samples) - [Build a minimal MCP client](build-mcp-client.md) - [Publish a package](/nuget/nuget-org/publish-a-package) - [Find and evaluate NuGet packages for your project](/nuget/consume-packages/finding-and-choosing-packages) - [What's new in .NET 10](../../core/whats-new/dotnet-10/overview.md) -
build-vector-search-app.md 9.5 KB
--- title: Quickstart - Build a minimal .NET AI RAG app description: Create an AI powered app to search and integrate with vector stores using embeddings and the Microsoft.Extensions.VectorData package for .NET ms.date: 03/04/2026 ms.topic: quickstart zone_pivot_groups: openai-library ai-usage: ai-assisted --- # Build a .NET AI vector search app In this quickstart, you create a .NET console app to perform semantic search on a _vector store_ to find relevant results for the user's query. You learn how to generate embeddings for user prompts and use those embeddings to query the vector data store. Vector stores, or vector databases, are essential for tasks like semantic search, retrieval augmented generation (RAG), and other scenarios that require grounding generative AI responses. While relational databases and document databases are optimized for structured and semi-structured data, vector databases are built to efficiently store, index, and manage data represented as embedding vectors. As a result, the indexing and search algorithms used by vector databases are optimized to efficiently retrieve data that can be used downstream in your applications. ## About the libraries The app uses the <xref:Microsoft.Extensions.AI> and <xref:Microsoft.Extensions.VectorData> libraries so you can write code using AI abstractions rather than a specific SDK. AI abstractions help create loosely coupled code that allows you to change the underlying AI model with minimal app changes. [📦 Microsoft.Extensions.VectorData.Abstractions](https://www.nuget.org/packages/Microsoft.Extensions.VectorData.Abstractions/) is a .NET library that provides a unified layer of abstractions for interacting with vector stores. The abstractions in `Microsoft.Extensions.VectorData.Abstractions` provide library authors and developers with the following functionality: - Perform create-read-update-delete (CRUD) operations on vector stores. - Use vector and text search on vector stores. <!--Prerequisites section--> :::zone target="docs" pivot="openai" [!INCLUDE [openai-prereqs](includes/prerequisites-openai.md)] :::zone-end :::zone target="docs" pivot="azure-openai" [!INCLUDE [azure-openai-prereqs](includes/prerequisites-azure-openai.md)] :::zone-end ## Create the app Complete the following steps to create a .NET console app that can: - Create and populate a vector store by generating embeddings for a data set. - Generate an embedding for the user prompt. - Query the vector store using the user prompt embedding. - Display the relevant results from the vector search. 1. In an empty directory on your computer, use the `dotnet new` command to create a new console app: ```dotnetcli dotnet new console -o VectorDataAI ``` 1. Change directory into the app folder: ```dotnetcli cd VectorDataAI ``` 1. Install the required packages: :::zone target="docs" pivot="azure-openai" ```bash dotnet add package Azure.Identity dotnet add package Azure.AI.OpenAI dotnet add package Microsoft.Extensions.AI.OpenAI dotnet add package Microsoft.Extensions.VectorData.Abstractions dotnet add package Microsoft.SemanticKernel.Connectors.InMemory --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets dotnet add package System.Linq.AsyncEnumerable ``` The following list describes each package in the `VectorDataAI` app: - [`Azure.Identity`](https://www.nuget.org/packages/Azure.Identity) provides [`Microsoft Entra ID`](/entra/fundamentals/whatis) token authentication support across the Azure SDK using classes such as `DefaultAzureCredential`. - [`Azure.AI.OpenAI`](https://www.nuget.org/packages/Azure.AI.OpenAI) is the official package for using OpenAI's .NET library with the Azure OpenAI Service. - [`Microsoft.Extensions.VectorData.Abstractions`](https://www.nuget.org/packages/Microsoft.Extensions.VectorData.Abstractions) enables Create-Read-Update-Delete (CRUD) and search operations on vector stores. - [`Microsoft.SemanticKernel.Connectors.InMemory`](https://www.nuget.org/packages/Microsoft.SemanticKernel.Connectors.InMemory) provides an in-memory vector store class to hold queryable vector data records. - [Microsoft.Extensions.Configuration](https://www.nuget.org/packages/Microsoft.Extensions.Configuration) provides an implementation of key-value pair—based configuration. - [`Microsoft.Extensions.Configuration.UserSecrets`](https://www.nuget.org/packages/Microsoft.Extensions.Configuration.UserSecrets) is a user secrets configuration provider implementation for `Microsoft.Extensions.Configuration`. :::zone-end :::zone target="docs" pivot="openai" ```bash dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.VectorData.Abstractions dotnet add package Microsoft.SemanticKernel.Connectors.InMemory --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets dotnet add package System.Linq.AsyncEnumerable ``` The following list describes each package in the `VectorDataAI` app: - [`Microsoft.Extensions.AI.OpenAI`](https://www.nuget.org/packages/Microsoft.Extensions.AI.OpenAI) provides AI abstractions for OpenAI-compatible models or endpoints. This library also includes the official [`OpenAI`](https://www.nuget.org/packages/OpenAI) library for the OpenAI service API as a dependency. - [`Microsoft.Extensions.VectorData.Abstractions`](https://www.nuget.org/packages/Microsoft.Extensions.VectorData.Abstractions) enables Create-Read-Update-Delete (CRUD) and search operations on vector stores. - [`Microsoft.SemanticKernel.Connectors.InMemory`](https://www.nuget.org/packages/Microsoft.SemanticKernel.Connectors.InMemory) provides an in-memory vector store class to hold queryable vector data records. - [Microsoft.Extensions.Configuration](https://www.nuget.org/packages/Microsoft.Extensions.Configuration) provides an implementation of key-value pair—based configuration. - [`Microsoft.Extensions.Configuration.UserSecrets`](https://www.nuget.org/packages/Microsoft.Extensions.Configuration.UserSecrets) is a user secrets configuration provider implementation for `Microsoft.Extensions.Configuration`. :::zone-end 1. Open the app in Visual Studio Code (or your editor of choice). ```bash code . ``` :::zone target="docs" pivot="azure-openai" [!INCLUDE [create-ai-service](includes/create-ai-service.md)] :::zone-end :::zone target="docs" pivot="openai" ## Configure the app 1. Navigate to the root of your .NET project from a terminal or command prompt. 1. Run the following commands to configure your OpenAI API key as a secret for the sample app: ```bash dotnet user-secrets init dotnet user-secrets set OpenAIKey <your-OpenAI-key> dotnet user-secrets set ModelName <your-OpenAI-model-name> ``` :::zone-end > [!NOTE] > For the model name, you need to specify a text embedding model such as `text-embedding-3-small` or `text-embedding-3-large` to generate embeddings for vector search in the sections that follow. For more information about embedding models, see [Embeddings](/azure/ai-services/openai/concepts/models#embeddings). ## Add the app code 1. Add a new class named `CloudService` to your project with the following properties: The <xref:Microsoft.Extensions.VectorData> attributes, such as <xref:Microsoft.Extensions.VectorData.VectorStoreKeyAttribute>, influence how each property is handled when used in a vector store. The `Vector` property stores a generated embedding that represents the semantic meaning of the `Description` value for vector searches. 1. In the `Program.cs` file, add the following code to create a data set that describes a collection of cloud services: 1. Create and configure an `IEmbeddingGenerator` implementation to send requests to an embedding AI model: :::zone target="docs" pivot="azure-openai" > [!NOTE] > <xref:Azure.Identity.DefaultAzureCredential> searches for authentication credentials from your local tooling. You'll need to assign the `Azure AI Developer` role to the account you used to sign in to Visual Studio or the Azure CLI. For more information, see [Authenticate to Foundry tools with .NET](../azure-ai-services-authentication.md). :::zone-end :::zone target="docs" pivot="openai" :::zone-end 1. Create and populate a vector store with the cloud service data. Use the `IEmbeddingGenerator` implementation to create and assign an embedding vector for each record in the cloud service data: The embeddings are numerical representations of the semantic meaning for each data record, which makes them compatible with vector search features. 1. Create an embedding for a search query and use it to perform a vector search on the vector store: 1. Use the `dotnet run` command to run the app: ```dotnetcli dotnet run ``` The app prints out the top result of the vector search, which is the cloud service that's most relevant to the original query. You can modify the query to try different search scenarios. :::zone target="docs" pivot="azure-openai" ## Clean up resources If you no longer need them, delete the Azure OpenAI resource and model deployment. 1. In the [Azure portal](https://aka.ms/azureportal), navigate to the Azure OpenAI resource. 1. Select the Azure OpenAI resource, and then select **Delete**. :::zone-end ## Next steps - [Quickstart - Chat with a local AI model](chat-local-model.md) - [Generate images from text using AI](text-to-image.md) -
chat-local-model.md 6.3 KB
--- title: Quickstart - Connect to and chat with a local AI using .NET description: Set up a local AI model and chat with it using a .NET console app and the Microsoft.Extensions.AI libraries ms.date: 03/04/2026 ms.topic: quickstart ai-usage: ai-assisted --- # Chat with a local AI model using .NET In this quickstart, you learn how to create a conversational .NET console chat app using an OpenAI or Azure OpenAI model. The app uses the <xref:Microsoft.Extensions.AI> library so you can write code using AI abstractions rather than a specific SDK. AI abstractions enable you to change the underlying AI model with minimal code changes. ## Prerequisites * [Install .NET 8.0](https://dotnet.microsoft.com/download) or higher * [Install Ollama](https://ollama.com/) locally on your device * [Visual Studio Code](https://code.visualstudio.com/) (optional) ## Run the local AI model Complete the following steps to configure and run a local AI model on your device. Many different AI models are available to run locally and are trained for different tasks, such as generating code, analyzing images, generative chat, or creating embeddings. For this quickstart, you'll use the general purpose `phi3:mini` model, which is a small but capable generative AI created by Microsoft. 1. Open a terminal window and verify that Ollama is available on your device: ```bash ollama ``` If Ollama is available, it displays a list of available commands. 1. Start Ollama: ```bash ollama serve ``` If Ollama is running, it displays a list of available commands. 1. Pull the `phi3:mini` model from the Ollama registry and wait for it to download: ```bash ollama pull phi3:mini ``` 1. After the download completes, run the model: ```bash ollama run phi3:mini ``` Ollama starts the `phi3:mini` model and provides a prompt for you to interact with it. ## Create the .NET app Complete the following steps to create a .NET console app that connects to your local `phi3:mini` AI model. 1. In a terminal window, navigate to an empty directory on your device and create a new app with the `dotnet new` command: ```dotnetcli dotnet new console -o LocalAI ``` 1. Change directory into the app folder: ```dotnetcli cd LocalAI ``` 1. Add the [OllamaSharp](https://www.nuget.org/packages/OllamaSharp) package to your app: ```dotnetcli dotnet add package OllamaSharp ``` 1. Open the new app in your editor of choice, such as Visual Studio Code. ```dotnetcli code . ``` ## Connect to and chat with the AI model In the steps ahead, you'll create a simple app that connects to the local AI and stores conversation history to improve the chat experience. 1. Open the _Program.cs_ file and replace the contents of the file with the following code: The preceding code accomplishes the following: * Creates an `OllamaChatClient` that implements the `IChatClient` interface. * This interface provides a loosely coupled abstraction you can use to chat with AI Models. * You can later change the underlying chat client implementation to another model, such as Azure OpenAI, without changing any other code. * Creates a `ChatHistory` object to store the messages between the user and the AI model. * Retrieves a prompt from the user and stores it in the `ChatHistory`. * Sends the chat data to the AI model to generate a response. > [!NOTE] > Ollama runs on port 11434 by default, which is why the AI model endpoint is set to `http://localhost:11434`. 1. Run the app and enter a prompt into the console to receive a response from the AI, such as the following: ```output Your prompt: Tell me three facts about .NET. AI response: 1. **Cross-Platform Development:** One of the significant strengths of .NET, particularly its newer iterations (.NET Core and .NET 5+), is cross-platform support. It allows developers to build applications that run on Windows, Linux, macOS, and various other operating systems seamlessly, enhancing flexibility and reducing barriers for a wider range of users. 2. **Rich Ecosystem and Library Support:** .NET has a rich ecosystem, comprising an extensive collection of libraries (such as those provided by the official NuGet Package Manager), tools, and services. This allows developers to work on web applications (.NET for desktop apps and ASP.NET Core for modern web applications), mobile applications (.NET MAUI), IoT solutions, AI/ML projects, and much more with a vast array of prebuilt components available at their disposal. 3. **Type Safety:** .NET operates under the Common Language Infrastructure (CLI) model and employs managed code for executing applications. This approach inherently offers strong type safety checks which help in preventing many runtime errors that are common in languages like C/C++. It also enables features such as garbage collection, thus relieving developers from manual memory management. These characteristics enhance the reliability of .NET-developed software and improve productivity by catching issues early during development. ``` 1. The response from the AI is accurate, but also verbose. The stored chat history enables the AI to modify its response. Instruct the AI to shorten the list it provided: ```output Your prompt: Shorten the length of each item in the previous response. AI Response: **Cross-platform Capabilities:** .NET allows building for various operating systems through platforms like .NET Core, promoting accessibility (Windows, Linux, macOS). **Extensive Ecosystem:** Offers a vast library selection via NuGet and tools for web (.NET Framework), mobile development (.NET MAUI), IoT, AI, providing rich capabilities to developers. **Type Safety & Reliability:** .NET's CLI model enforces strong typing and automatic garbage collection, mitigating runtime errors, thus enhancing application stability. ``` The updated response from the AI is much shorter the second time. Due to the available chat history, the AI was able to assess the previous result and provide shorter summaries. ## Next steps * [Generate text and conversations with .NET and Azure OpenAI Completions](/training/modules/open-ai-dotnet-text-completions/) -
create-assistant.md 5.2 KB
--- title: Quickstart - Create a minimal AI assistant using .NET description: Learn to create a minimal AI assistant with tooling capabilities using .NET and the Azure OpenAI SDK libraries ms.date: 03/04/2026 ms.topic: quickstart zone_pivot_groups: openai-library ai-usage: ai-assisted --- # Create a minimal AI assistant using .NET In this quickstart, you'll learn how to create a minimal AI assistant using the OpenAI or Azure OpenAI SDK libraries. AI assistants provide agentic functionality to help users complete tasks using AI tools and models. In the sections ahead, you'll learn the following: - Core components and concepts of AI assistants - How to create an assistant using the Azure OpenAI SDK - How to enhance and customize the capabilities of an assistant ## Prerequisites ::: zone pivot="openai" * [Install .NET 8.0](https://dotnet.microsoft.com/download) or higher * [Visual Studio Code](https://code.visualstudio.com/) (optional) * [Visual Studio](https://visualstudio.com/) (optional) * An access key for an OpenAI model :::zone-end ::: zone pivot="azure-openai" * [Install .NET 8.0](https://dotnet.microsoft.com/download) or higher * [Visual Studio Code](https://code.visualstudio.com/) (optional) * [Visual Studio](https://visualstudio.com/) (optional) * Access to an Azure OpenAI instance via Azure Identity or an access key :::zone-end ## Core components of AI assistants AI assistants are based around conversational threads with a user. The user sends prompts to the assistant on a conversation thread, which direct the assistant to complete tasks using the tools it has available. Assistants can process and analyze data, make decisions, and interact with users or other systems to achieve specific goals. Most assistants include the following components: | **Component** | **Description** | |---------------|-----------------| | **Assistant** | The core AI client and logic that uses Azure OpenAI models, manages conversation threads, and utilizes configured tools. | | **Thread** | A conversation session between an assistant and a user. Threads store messages and automatically handle truncation to fit content into a model's context. | | **Message** | A message created by an assistant or a user. Messages can include text, images, and other files. Messages are stored as a list on the thread. | | **Run** | Activation of an assistant to begin running based on the contents of the thread. The assistant uses its configuration and the thread's messages to perform tasks by calling models and tools. As part of a run, the assistant appends messages to the thread. | | **Run steps** | A detailed list of steps the assistant took as part of a run. An assistant can call tools or create messages during its run. Examining run steps allows you to understand how the assistant is getting to its final results. | Assistants can also be configured to use multiple tools in parallel to complete tasks, including the following: - **Code interpreter tool**: Writes and runs code in a sandboxed execution environment. - **Function calling**: Runs local custom functions you define in your code. - **File search capabilities**: Augments the assistant with knowledge from outside its model. By understanding these core components and how they interact, you can build and customize powerful AI assistants to meet your specific needs. ## Create the .NET app Complete the following steps to create a .NET console app and add the package needed to work with assistants: ::: zone pivot="openai" 1. In a terminal window, navigate to an empty directory on your device and create a new app with the `dotnet new` command: ```dotnetcli dotnet new console -o AIAssistant ``` 1. Add the [OpenAI](https://www.nuget.org/packages/OpenAI) package to your app: ```dotnetcli dotnet add package OpenAI ``` 1. Open the new app in your editor of choice, such as Visual Studio Code. ```dotnetcli code . ``` ::: zone-end ::: zone pivot="azure-openai" 1. In a terminal window, navigate to an empty directory on your device and create a new app with the `dotnet new` command: ```dotnetcli dotnet new console -o AIAssistant ``` 1. Add the [Azure.AI.OpenAI](https://www.nuget.org/packages/Azure.AI.OpenAI) package to your app: ```dotnetcli dotnet add package Azure.AI.OpenAI ``` 1. Open the new app in your editor of choice, such as Visual Studio Code. ```dotnetcli code . ``` ::: zone-end ## Create the AI assistant client 1. Open the `Program.cs` file and replace the contents of the file with the following code to create the required clients: 1. Create an in-memory sample document and upload it to the `OpenAIFileClient`: 1. Enable file search and code interpreter tooling capabilities via the `AssistantCreationOptions`: 1. Create the `Assistant` and a thread to manage interactions between the user and the assistant: 1. Print the messages and save the generated image from the conversation with the assistant: Locate and open the saved image in the app `bin` directory, which should resemble the following: ## Next steps - [Generate text and conversations with .NET and Azure OpenAI Completions](/training/modules/open-ai-dotnet-text-completions/) -
generate-images.md 4 KB
--- title: Quickstart - Generate images using OpenAI.Images.ImageClient description: Create a simple app using to generate images using OpenAI.Images.ImageClient in .NET. ms.date: 03/04/2026 ms.topic: quickstart zone_pivot_groups: openai-library ai-usage: ai-assisted --- # Generate images using OpenAI.Images.ImageClient In this quickstart, you create a .NET console app that uses `OpenAI.Images.ImageClient` to generate images using an OpenAI or Azure OpenAI DALL-E AI model. These models generate images from text prompts. :::zone target="docs" pivot="openai" [!INCLUDE [openai-prereqs](includes/prerequisites-openai.md)] :::zone-end :::zone target="docs" pivot="azure-openai" [!INCLUDE [azure-openai-prereqs](includes/prerequisites-azure-openai.md)] :::zone-end ## Create the app Complete the following steps to create a .NET console app to connect to an AI model. 1. In an empty directory on your computer, use the `dotnet new` command to create a new console app: ```dotnetcli dotnet new console -o ImagesAI ``` 1. Change directory into the app folder: ```dotnetcli cd ImagesAI ``` 1. Install the required packages: :::zone target="docs" pivot="azure-openai" ```bash dotnet add package Azure.AI.OpenAI dotnet add package Azure.Identity dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` :::zone-end :::zone target="docs" pivot="openai" ```bash dotnet add package OpenAI dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` :::zone-end 1. Open the app in Visual Studio Code or your editor of choice. ```bash code . ``` :::zone target="docs" pivot="azure-openai" [!INCLUDE [create-ai-service](includes/create-ai-service.md)] :::zone-end :::zone target="docs" pivot="openai" ## Configure the app 1. Navigate to the root of your .NET project from a terminal or command prompt. 1. Run the following commands to configure your OpenAI API key as a secret for the sample app: ```bash dotnet user-secrets init dotnet user-secrets set OpenAIKey <your-OpenAI-key> dotnet user-secrets set ModelName <your-OpenAI-model-name> ``` :::zone-end ## Add the app code 1. In the `Program.cs` file, add the following code to connect and authenticate to the AI model. :::zone target="docs" pivot="azure-openai" > [!NOTE] > <xref:Azure.Identity.DefaultAzureCredential> searches for authentication credentials from your local tooling. If you aren't using the `azd` template to provision the Azure OpenAI resource, you'll need to assign the `Azure AI Developer` role to the account you used to sign in to Visual Studio or the Azure CLI. For more information, see [Authenticate to Foundry tools with .NET](../azure-ai-services-authentication.md). :::zone-end :::zone target="docs" pivot="openai" :::zone-end The preceding code: - Reads essential configuration values from the project user secrets to connect to the AI model. - Creates an `OpenAI.Images.ImageClient` to connect to the AI model. - Sends a prompt to the model that describes the desired image. - Prints the URL of the generated image to the console output. 1. Run the app: ```dotnetcli dotnet run ``` Navigate to the image URL in the console output to view the generated image. Customize the text content of the prompt to create new images or modify the original. :::zone target="docs" pivot="azure-openai" ## Clean up resources If you no longer need them, delete the Azure OpenAI resource and GPT-4 model deployment. 1. In the [Azure portal](https://aka.ms/azureportal), navigate to the Azure OpenAI resource. 1. Select the Azure OpenAI resource, and then select **Delete**. :::zone-end ## Next steps - [Quickstart - Build an AI chat app with .NET](build-chat-app.md) - [Generate text and conversations with .NET and Azure OpenAI Completions](/training/modules/open-ai-dotnet-text-completions/) -
process-data.md 7.1 KB
--- title: Quickstart - Process custom data for AI description: Create a data ingestion pipeline to process and prepare custom data for AI applications using Microsoft.Extensions.DataIngestion. ms.date: 12/11/2025 ms.topic: quickstart ai-usage: ai-assisted --- # Process custom data for AI applications In this quickstart, you learn how to create a data ingestion pipeline to process and prepare custom data for AI applications. The app uses the <xref:Microsoft.Extensions.DataIngestion> library to read documents, enrich content with AI, chunk text semantically, and store embeddings in a vector database for semantic search. Data ingestion is essential for retrieval-augmented generation (RAG) scenarios where you need to process large amounts of unstructured data and make it searchable for AI applications. [!INCLUDE [azure-openai-prereqs](includes/prerequisites-azure-openai.md)] ## Create the app Complete the following steps to create a .NET console app. 1. In an empty directory on your computer, use the `dotnet new` command to create a new console app: ```dotnetcli dotnet new console -o ProcessDataAI ``` 1. Change directory into the app folder: ```dotnetcli cd ProcessDataAI ``` 1. Install the required packages: ```bash dotnet add package Azure.AI.OpenAI dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets dotnet add package Microsoft.Extensions.DataIngestion --prerelease dotnet add package Microsoft.Extensions.DataIngestion.Markdig --prerelease dotnet add package Microsoft.Extensions.Logging.Console dotnet add package Microsoft.ML.Tokenizers.Data.O200kBase dotnet add package Microsoft.SemanticKernel.Connectors.SqliteVec --prerelease ``` ## Create the AI service 1. To provision an Azure OpenAI service and model, complete the steps in the [Create and deploy an Azure OpenAI Service resource](/azure/ai-services/openai/how-to/create-resource) article. For this quickstart, you need to provision two models: `gpt-5` and `text-embedding-3-small`. 1. From a terminal or command prompt, navigate to the root of your project directory. 1. Run the following commands to configure your Azure OpenAI endpoint and API key for the sample app: ```bash dotnet user-secrets init dotnet user-secrets set AZURE_OPENAI_ENDPOINT <your-Azure-OpenAI-endpoint> dotnet user-secrets set AZURE_OPENAI_API_KEY <your-Azure-OpenAI-API-key> ``` ## Open the app in an editor Open the app in Visual Studio Code (or your editor of choice). ```bash code . ``` ## Create the sample data 1. Copy the [sample.md](https://raw.githubusercontent.com/dotnet/docs/refs/heads/main/docs/ai/quickstarts/snippets/process-data/data/sample.md) file to a folder named `data` in your project directory. 1. Configure the project to copy this file to the output directory. If you're using Visual Studio, right-click on the file in Solution Explorer, select **Properties**, and then set **Copy to Output Directory** to **Copy if newer**. ## Add the app code The data ingestion pipeline consists of several components that work together to process documents: - **Document reader**: Reads Markdown files from a directory. - **Document processor**: Enriches images with AI-generated alternative text. - **Chunker**: Splits documents into semantic chunks using embeddings. - **Chunk processor**: Generates AI summaries for each chunk. - **Vector store writer**: Stores chunks with embeddings in a SQLite database. 1. In the `Program.cs` file, delete any existing code and add the following code to configure the document reader: The <xref:Microsoft.Extensions.DataIngestion.MarkdownReader> class reads Markdown documents and converts them into a unified format that works well with large language models. 1. Add code to configure logging for the pipeline: 1. Add code to configure the AI client for enrichment and chat: 1. Add code to configure the document processor that enriches images with AI-generated descriptions: The <xref:Microsoft.Extensions.DataIngestion.ImageAlternativeTextEnricher> uses large language models to generate descriptive alternative text for images within documents. That text makes them more accessible and improves their semantic meaning. 1. Add code to configure the embedding generator for creating vector representations: [Embeddings](../conceptual/embeddings.md) are numerical representations of the semantic meaning of text, which enables vector similarity search. 1. Add code to configure the chunker that splits documents into semantic chunks: The <xref:Microsoft.Extensions.DataIngestion.Chunkers.SemanticSimilarityChunker> intelligently splits documents by analyzing the semantic similarity between sentences, ensuring that related content stays together. This process produces chunks that preserve meaning and context better than simple character or token-based chunking. 1. Add code to configure the chunk processor that generates summaries: The <xref:Microsoft.Extensions.DataIngestion.SummaryEnricher> automatically generates concise summaries for each chunk, which can improve retrieval accuracy by providing a high-level overview of the content. 1. Add code to configure the SQLite vector store for storing embeddings: The vector store stores chunks along with their embeddings, enabling fast semantic search capabilities. 1. Add code to compose all the components into a complete pipeline: The <xref:Microsoft.Extensions.DataIngestion.IngestionPipeline`1> combines all the components into a cohesive workflow that processes documents from start to finish. 1. Add code to process documents from a directory: The pipeline processes all Markdown files in the `./data` directory and reports the status of each document. 1. Add code to enable interactive search of the processed documents: The search functionality converts user queries into embeddings and finds the most semantically similar chunks in the vector store. ## Run the app 1. Use the `dotnet run` command to run the app: ```dotnetcli dotnet run ``` The app processes all Markdown files in the `./data` directory and displays the processing status for each document. Once processing is complete, you can enter natural language questions to search the processed content. 1. Enter a question at the prompt to search the data: ```output Enter your question (or 'exit' to quit): What is data ingestion? ``` The app returns the most relevant chunks from your documents along with their similarity scores. 1. Type `exit` to quit the application. ## Clean up resources If you no longer need them, delete the Azure OpenAI resource and model deployment. 1. In the [Azure Portal](https://aka.ms/azureportal), navigate to the Azure OpenAI resource. 1. Select the Azure OpenAI resource, and then select **Delete**. ## Next steps - [Data ingestion concepts](../conceptual/data-ingestion.md) - [Implement RAG using vector search](../tutorials/tutorial-ai-vector-search.md) - [Build a .NET AI vector search app](build-vector-search-app.md) -
prompt-model.md 4.7 KB
--- title: Quickstart - Connect to and prompt an AI model with .NET description: Create a simple chat app using Microsoft.Extensions.AI to summarize a text. ms.date: 03/04/2026 ms.topic: quickstart zone_pivot_groups: openai-library ai-usage: ai-assisted --- # Connect to and prompt an AI model In this quickstart, you learn how to create a .NET console chat app to connect to and prompt an OpenAI or Azure OpenAI model. The app uses the <xref:Microsoft.Extensions.AI> library so you can write code using AI abstractions rather than a specific SDK. AI abstractions enable you to change the underlying AI model with minimal code changes. :::zone target="docs" pivot="openai" [!INCLUDE [openai-prereqs](includes/prerequisites-openai.md)] :::zone-end :::zone target="docs" pivot="azure-openai" [!INCLUDE [azure-openai-prereqs](includes/prerequisites-azure-openai.md)] :::zone-end ## Create the app Complete the following steps to create a .NET console app to connect to an AI model. 1. In an empty directory on your computer, use the `dotnet new` command to create a new console app: ```dotnetcli dotnet new console -o ExtensionsAI ``` 1. Change directory into the app folder: ```dotnetcli cd ExtensionsAI ``` 1. Install the required packages: :::zone target="docs" pivot="azure-openai" ```bash dotnet add package Azure.AI.OpenAI dotnet add package Azure.Identity dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` :::zone-end :::zone target="docs" pivot="openai" ```bash dotnet add package OpenAI dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` :::zone-end 1. Open the app in Visual Studio Code or your editor of choice. :::zone target="docs" pivot="azure-openai" [!INCLUDE [create-ai-service](includes/create-ai-service.md)] :::zone-end :::zone target="docs" pivot="openai" ## Configure the app 1. Navigate to the root of your .NET project from a terminal or command prompt. 1. Run the following commands to configure your OpenAI API key as a secret for the sample app: ```bash dotnet user-secrets init dotnet user-secrets set OpenAIKey <your-OpenAI-key> dotnet user-secrets set ModelName <your-OpenAI-model-name> ``` :::zone-end ## Add the app code The app uses the [`Microsoft.Extensions.AI`](https://www.nuget.org/packages/Microsoft.Extensions.AI/) package to send and receive requests to the AI model. 1. Copy the [benefits.md](https://raw.githubusercontent.com/dotnet/docs/refs/heads/main/docs/ai/quickstarts/snippets/prompt-completion/azure-openai/benefits.md) file to your project directory. Configure the project to copy this file to the output directory. If you're using Visual Studio, right-click on the file in Solution Explorer, select **Properties**, and then set **Copy to Output Directory** to **Copy if newer**. 1. In the `Program.cs` file, add the following code to connect and authenticate to the AI model. :::zone target="docs" pivot="azure-openai" > [!NOTE] > <xref:Azure.Identity.DefaultAzureCredential> searches for authentication credentials from your local tooling. If you aren't using the `azd` template to provision the Azure OpenAI resource, you'll need to assign the `Azure AI Developer` role to the account you used to sign in to Visual Studio or the Azure CLI. For more information, see [Authenticate to Foundry tools with .NET](../azure-ai-services-authentication.md). :::zone-end :::zone target="docs" pivot="openai" :::zone-end 1. Add code to read the `benefits.md` file content and then create a prompt for the model. The prompt instructs the model to summarize the file's text content in 20 words or less. 1. Call the `GetResponseAsync` method to send the prompt to the model to generate a response. 1. Run the app: ```dotnetcli dotnet run ``` The app prints out the completion response from the AI model. Customize the text content of the `benefits.md` file or the length of the summary to see the differences in the responses. :::zone target="docs" pivot="azure-openai" ## Clean up resources If you no longer need them, delete the Azure OpenAI resource and GPT-4 model deployment. 1. In the [Azure portal](https://aka.ms/azureportal), navigate to the Azure OpenAI resource. 1. Select the Azure OpenAI resource, and then select **Delete**. :::zone-end ## Next steps - [Quickstart - Build an AI chat app with .NET](build-chat-app.md) - [Generate text and conversations with .NET and Azure OpenAI Completions](/training/modules/open-ai-dotnet-text-completions/) -
publish-mcp-registry.md 12.2 KB
--- title: Quickstart - Publish a .NET MCP server to the MCP Registry description: Learn how to publish your NuGet-based MCP server to the Official MCP Registry, including creating a server.json manifest, updating your package README, and using the MCP Publisher tool. ms.date: 03/04/2026 ms.topic: quickstart author: joelverhagen zone_pivot_groups: operating-systems-set-one ai-usage: ai-assisted --- # Publish an MCP server on NuGet.org to the Official MCP Registry In this quickstart, you publish your NuGet-based local MCP server to the [Official MCP Registry](https://github.com/modelcontextprotocol/registry/blob/main/docs/design/ecosystem-vision.md). The Official MCP Registry is an *upstream data source* for the MCP ecosystem. Other MCP registries, such as the [GitHub MCP Registry](https://github.com/mcp), will soon use the Official MCP Registry as a source of MCP server listings. > [!NOTE] > This guide focuses on publishing **local MCP servers** packaged with NuGet. The Official MCP Registry also supports **remote MCP servers**. While the `server.json` publishing process is similar for remote servers, their configuration requires a URL instead of a package manager reference. Remote servers can be implemented in any language. For an example, see [an Azure Functions code sample for a .NET remote MCP server](/samples/azure-samples/remote-mcp-functions-dotnet/remote-mcp-functions-dotnet/). ## Prerequisites - A [GitHub account](https://github.com/join) - [Visual Studio Code](https://code.visualstudio.com/) - Your MCP server is packaged with NuGet and published to NuGet.org ([quickstart](./build-mcp-server.md)). ## Create a server.json manifest file *If you used the NuGet MCP server quickstart and `mcpserver` project template, you can skip this step.* 1. Navigate to your MCP server's source directory and create a new `server.json` file. ::: zone pivot="os-windows" ```powershell cd path\my\project # create and open the server.json file code .mcp\server.json ``` ::: zone-end ::: zone pivot="os-linux" ```bash cd path/my/project # create and open the server.json file code .mcp/server.json ``` ::: zone-end ::: zone pivot="os-macos" ```bash cd path/my/project # create and open the server.json file code .mcp/server.json ``` ::: zone-end 2. Use this content to start with and fill in the placeholders. 3. Save the file. Use this reference to understand more about the fields: | Property | Example | Purpose | | ----------------------- | -------------------------------------- | ----------------------------------------------------------------------------------------------------------- | | `name` | `io.github.contoso/data-mcp` | Unique identifier for the MCP server, namespaced using reverse DNS names, **case sensitive** | | `version` | `0.1.0-beta` | Version of the MCP server listing<br>Consider using the same version as the MCP server package on NuGet.org | | `description` | `Access Contoso data in your AI agent` | Description of your MCP server, up to 100 characters | | `title` | `Contoso Data` | Optional: short human-readable title, up to 100 characters | | `websiteUrl` | `https://contoso.com/docs/mcp` | Optional: URL to the server's homepage, documentation, or project website | | `packages` `identifier` | `Contoso.Data.Mcp` | The ID of your MCP server package on NuGet.org | | `packages` `version` | `0.1.0-beta` | The version of your MCP server package on NuGet.org | | `repository` `url` | `https://github.com/contoso/data-mcp` | Optional: GitHub repository URL | The `name` field has two parts, separated by a forward slash `/`. The first part is a namespace based off of a reverse DNS name. The authentication method you use in later steps will give you access to a specific namespace. For example, using GitHub-based authentication will give you access to `io.github.<your GitHub username>/*`. The second part, after the forward slash, is a custom identifier for your server within the namespace. Think of this much like a NuGet package ID. It should be unchanging and descriptive of your MCP server. Using your GitHub repository name is a reasonable option if you only have one MCP server published from that repository. ## Update your package README The Official MCP Registry verifies that your MCP server package references the `name` specified in your `server.json` file. 1. If you haven't already, add a README.md to your MCP server NuGet package. See [how to do this in your project file](/nuget/reference/msbuild-targets#packagereadmefile). 2. Open the README.md used by your NuGet package. ::: zone pivot="os-windows" ```powershell code path\to\README.md ``` ::: zone-end ::: zone pivot="os-linux" ```bash code path/to/README.md ``` ::: zone-end ::: zone pivot="os-macos" ```bash code path/to/README.md ``` ::: zone-end 3. Add the following line to your README.md. Since it is enclosed in an HTML comment, it can be anywhere and won't be rendered. ```markdown <!-- mcp-name: [name property from the server.json] --> ``` Example: ```markdown <!-- mcp-name: io.github.contoso/data-mcp --> ``` 4. Save the README.md file. ## Publish your MCP server package to NuGet.org Because your README.md now has an `mcp-name` declared in it, publish the latest package to NuGet.org. 1. If needed, update your package version and the respective version strings in your `server.json`. 2. Pack your project so the latest README.md version is contained. ```bash dotnet pack ``` 3. Push it to NuGet.org either [via the website](https://www.nuget.org/packages/manage/upload) or using the CLI: ::: zone pivot="os-windows" ```powershell dotnet push bin\Release\*.nupkg -k <API key here> -s https://api.nuget.org/v3/index.json ``` ::: zone-end ::: zone pivot="os-linux" ```bash dotnet push bin/Release/*.nupkg -k <API key here> -s https://api.nuget.org/v3/index.json ``` ::: zone-end ::: zone pivot="os-macos" ```bash dotnet push bin/Release/*.nupkg -k <API key here> -s https://api.nuget.org/v3/index.json ``` ::: zone-end ## Wait for your package to become available NuGet.org performs validations against your package before making it available so you must wait to publish your MCP server to the Official MCP Registry since it verifies that the package is accessible. To wait for your package to become available, either continue to periodically refresh the package details page on NuGet.org until the validating message disappears, or use the following PowerShell script to poll for availability. ```powershell $id = "<your NuGet package ID here>".ToLowerInvariant() $version = "<your NuGet package version here>".ToLowerInvariant() $url = "https://api.nuget.org/v3-flatcontainer/$id/$version/readme" $elapsed = 0; $interval = 10; $timeout = 300 Write-Host "Checking for package README of $id $version." while ($true) { if ($elapsed -gt $timeout) { Write-Error "Package README is not available after $elapsed seconds. URL: $url" exit 1 } try { Invoke-WebRequest -Uri $url -ErrorAction Stop | Out-Null Write-Host "Package README is now available." break } catch { Write-Host "Package README is not yet available. Elapsed time: $elapsed seconds." Start-Sleep -Seconds $interval; $elapsed += $interval continue } } ``` This script can be leveraged in a CI/CD pipeline to ensure the next step (publishing to the Official MCP Registry) does not happen before the NuGet package is available. ## Download the MCP Publisher tool 1. Download the `mcp-publisher-*.tar.gz` file from the Official MCP Registry GitHub repository that matches your CPU architecture. ::: zone pivot="os-windows" - Windows x64: [mcp-publisher_windows_amd64.tar.gz](https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_windows_amd64.tar.gz) - Windows Arm64: [mcp-publisher_windows_arm64.tar.gz](https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_windows_arm64.tar.gz) ::: zone-end ::: zone pivot="os-linux" - Linux x64: [mcp-publisher_linux_amd64.tar.gz](https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_linux_amd64.tar.gz) - Linux Arm64: [mcp-publisher_linux_arm64.tar.gz](https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_linux_arm64.tar.gz) ::: zone-end ::: zone pivot="os-macos" - macOS x64: [mcp-publisher_darwin_amd64.tar.gz](https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_darwin_amd64.tar.gz) - macOS Arm64: [mcp-publisher_darwin_arm64.tar.gz](https://github.com/modelcontextprotocol/registry/releases/latest/download/mcp-publisher_darwin_arm64.tar.gz) ::: zone-end See the full list of assets in the [latest release](https://github.com/modelcontextprotocol/registry/releases/latest). 2. Extract the downloaded .tar.gz to the current directory. ::: zone pivot="os-windows" ```powershell # For Windows x64 tar xf 'mcp-publisher_windows_amd64.tar.gz' # For Windows ARM64 tar xf 'mcp-publisher_windows_arm64.tar.gz' ``` ::: zone-end ::: zone pivot="os-linux" ```bash # For Linux x64 tar xf 'mcp-publisher_linux_amd64.tar.gz' # For Linux ARM64 tar xf 'mcp-publisher_linux_arm64.tar.gz' ``` ::: zone-end ::: zone pivot="os-macos" ```bash # For macOS x64 tar xf 'mcp-publisher_darwin_amd64.tar.gz' # For macOS ARM64 tar xf 'mcp-publisher_darwin_arm64.tar.gz' ``` ::: zone-end ## Publish to the Official MCP Registry The Official MCP Registry has different authentication mechanisms based on the namespace MCP server's `name`. In this guide, we are using a namespace based on GitHub (`io.github.<your GitHub username>/*`) so GitHub authentication must be used. See the [registry documentation for information on other authentication modes](https://github.com/modelcontextprotocol/registry/blob/main/docs/modelcontextprotocol-io/authentication.mdx), which unlock other namespaces. 1. Log in using GitHub interactive authentication. ::: zone pivot="os-windows" ```powershell .\mcp-publisher.exe login github ``` ::: zone-end ::: zone pivot="os-linux" ```bash ./mcp-publisher login github ``` ::: zone-end ::: zone pivot="os-macos" ```bash ./mcp-publisher login github ``` ::: zone-end Follow the instructions provided by the tool. You will provide a code to GitHub in your web browser to complete the flow. Once the flow is complete, you will be able to publish `server.json` files to the `io.github.<your GitHub username>/*` namespace. 2. Publish your `server.json` file to the Official MCP Registry. ::: zone pivot="os-windows" ```powershell .\mcp-publisher.exe publish path\to\.mcp\server.json ``` ::: zone-end ::: zone pivot="os-linux" ```bash ./mcp-publisher publish path/to/.mcp/server.json ``` ::: zone-end ::: zone pivot="os-macos" ```bash ./mcp-publisher publish path/to/.mcp/server.json ``` ::: zone-end 3. When the command succeeds, you can verify that your MCP server is published by going to the [registry home page](https://registry.modelcontextprotocol.io/) and searching for your server name. ## Related content - [Build and publish an MCP server to NuGet.org](./build-mcp-server.md) - [Publish a NuGet package](/nuget/nuget-org/publish-a-package) - [Conceptual: MCP servers in NuGet Packages](/nuget/concepts/nuget-mcp) - [Get started with .NET AI and the Model Context Protocol](../get-started-mcp.md) -
structured-output.md 5.5 KB
--- title: Quickstart - Request a response with structured output description: Learn how to create a chat app that responds with structured output, that is, output that conforms to a type that you specify. ms.date: 03/04/2026 ms.topic: quickstart ai-usage: ai-assisted --- # Request a response with structured output In this quickstart, you create a chat app that requests a response with *structured output*. A structured output response is a chat response that's of a type you specify instead of just plain text. The chat app you create in this quickstart analyzes sentiment of various product reviews, categorizing each review according to the values of a custom enumeration. ## Prerequisites - [.NET 8 or a later version](https://dotnet.microsoft.com/download) - [Visual Studio Code](https://code.visualstudio.com/) (optional) ## Configure the AI service To provision an Azure OpenAI service and model using the Azure portal, complete the steps in the [Create and deploy an Azure OpenAI Service resource](/azure/ai-services/openai/how-to/create-resource?pivots=web-portal) article. In the "Deploy a model" step, select the `gpt-5` model. ## Create the chat app Complete the following steps to create a console app that connects to the `gpt-5` AI model. 1. In a terminal window, navigate to the directory where you want to create your app, and create a new console app with the `dotnet new` command: ```dotnetcli dotnet new console -o SOChat ``` 1. Navigate to the `SOChat` directory, and add the necessary packages to your app: ```dotnetcli dotnet add package Azure.AI.OpenAI dotnet add package Azure.Identity dotnet add package Microsoft.Extensions.AI dotnet add package Microsoft.Extensions.AI.OpenAI dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` 1. Run the following commands to add [app secrets](/aspnet/core/security/app-secrets) for your Azure OpenAI endpoint and tenant ID: ```bash dotnet user-secrets init dotnet user-secrets set AZURE_OPENAI_ENDPOINT <your-Azure-OpenAI-endpoint> dotnet user-secrets set AZURE_TENANT_ID <your-tenant-ID> ``` > [!NOTE] > Depending on your environment, the tenant ID might not be needed. In that case, remove it from the code that instantiates the <xref:Azure.Identity.DefaultAzureCredential>. 1. Open the new app in your editor of choice. ## Add the code 1. Define the enumeration that describes the different sentiments. 1. Create the <xref:Microsoft.Extensions.AI.IChatClient> that will communicate with the model. > [!NOTE] > <xref:Azure.Identity.DefaultAzureCredential> searches for authentication credentials from your environment or local tooling. You'll need to assign the `Azure AI Developer` role to the account you used to sign in to Visual Studio or the Azure CLI. For more information, see [Authenticate to Foundry tools with .NET](../azure-ai-services-authentication.md). 1. Send a request to the model with a single product review, and then print the analyzed sentiment to the console. You declare the requested structured output type by passing it as the type argument to the <xref:Microsoft.Extensions.AI.ChatClientStructuredOutputExtensions.GetResponseAsync``1(Microsoft.Extensions.AI.IChatClient,System.String,Microsoft.Extensions.AI.ChatOptions,System.Nullable{System.Boolean},System.Threading.CancellationToken)?displayProperty=nameWithType> extension method. This code produces output similar to: ```output Sentiment: Positive ``` 1. Instead of just analyzing a single review, you can analyze a collection of reviews. This code produces output similar to: ```output Review: Best purchase ever! | Sentiment: Positive Review: Returned it immediately. | Sentiment: Negative Review: Hello | Sentiment: Neutral Review: It works as advertised. | Sentiment: Neutral Review: The packaging was damaged but otherwise okay. | Sentiment: Neutral ``` 1. And instead of requesting just the analyzed enumeration value, you can request the text response along with the analyzed value. Define a [record type](../../csharp/language-reference/builtin-types/record.md) to contain the text response and analyzed sentiment: (This record type is defined using [primary constructor](../../csharp/programming-guide/classes-and-structs/instance-constructors.md#primary-constructors) syntax. Primary constructors combine the type definition with the parameters necessary to instantiate any instance of the class. The C# compiler generates public properties for the primary constructor parameters.) Send the request using the record type as the type argument to `GetResponseAsync<T>`: This code produces output similar to: ```output Response text: Certainly, I have analyzed the sentiment of the review you provided. Sentiment: Neutral ``` ## Clean up resources If you no longer need them, delete the Azure OpenAI resource and model deployment. 1. In the [Azure portal](https://aka.ms/azureportal), navigate to the Azure OpenAI resource. 1. Select the Azure OpenAI resource, and then select **Delete**. ## See also - [Structured outputs (Azure OpenAI Service)](/azure/ai-services/openai/how-to/structured-outputs) - [Using JSON schema for structured output in .NET for OpenAI models](https://devblogs.microsoft.com/semantic-kernel/using-json-schema-for-structured-output-in-net-for-openai-models) - [Introducing Structured Outputs in the API (OpenAI)](https://openai.com/index/introducing-structured-outputs-in-the-api/) -
text-to-image.md 10.5 KB
--- title: Quickstart - Generate images from text using AI description: Learn how to use Microsoft.Extensions.AI to generate images from text prompts using AI models in a .NET application. ms.date: 03/04/2026 ms.topic: quickstart ai-usage: ai-assisted --- # Generate images from text using AI In this quickstart, you use the <xref:Microsoft.Extensions.AI> (MEAI) library to generate images from text prompts using an AI model. The MEAI text-to-image capabilities let you generate images from natural language prompts or existing images using a consistent and extensible API surface. The <xref:Microsoft.Extensions.AI.IImageGenerator> interface provides a unified, extensible API for working with various image generation services, making it easy to integrate text-to-image capabilities into your .NET apps. The interface supports: - Text-to-image generation. - Pipeline composition with middleware (logging, telemetry, caching). - Flexible configuration options. - Support for multiple AI providers. > [!NOTE] > The `IImageGenerator` interface is currently marked as experimental with the `MEAI001` diagnostic ID. You might need to suppress this warning in your project file or code. <!--Prereqs--> [!INCLUDE [azure-openai-prereqs](../quickstarts/includes/prerequisites-azure-openai.md)] ## Configure the AI service To provision an Azure OpenAI service and model using the Azure portal, complete the steps in the [Create and deploy an Azure OpenAI Service resource](/azure/ai-services/openai/how-to/create-resource?pivots=web-portal) article. In the "Deploy a model" step, select the `gpt-image-1` model. > [!NOTE] > `gpt-image-1` is a newer model that offers several improvements over DALL-E 3. It's available from OpenAI on a limited basis; apply for access with [this form](https://aka.ms/oai/gptimage1access). ## Create the application Complete the following steps to create a .NET console application that generates images from text prompts. 1. Create a new console application: ```dotnetcli dotnet new console -o TextToImageAI ``` 1. Navigate to the `TextToImageAI` directory, and add the necessary packages to your app: ```dotnetcli dotnet add package Azure.AI.OpenAI dotnet add package Microsoft.Extensions.AI.OpenAI dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` 1. Run the following commands to add [app secrets](/aspnet/core/security/app-secrets) for your Azure OpenAI endpoint and API key: ```bash dotnet user-secrets init dotnet user-secrets set AZURE_OPENAI_ENDPOINT <your-Azure-OpenAI-endpoint> dotnet user-secrets set AZURE_OPENAI_API_KEY <your-azure-openai-api-key> ``` 1. Open the new app in your editor of choice (for example, Visual Studio). ## Implement basic image generation 1. Update the `Program.cs` file with the following code to get the configuration data and create the <xref:Azure.AI.OpenAI.AzureOpenAIClient>: The preceding code: - Loads configuration from user secrets. - Creates an `ImageClient` from the OpenAI SDK. - Converts the `ImageClient` to an `IImageGenerator` using the <xref:Microsoft.Extensions.AI.OpenAIClientExtensions.AsIImageGenerator(OpenAI.Images.ImageClient)> extension method. 1. Add the following code to implement basic text-to-image generation: The preceding code: - Sets the requested image file type by setting <xref:Microsoft.Extensions.AI.ImageGenerationOptions.MediaType?displayProperty=nameWithType>. - Generates an image using the <xref:Microsoft.Extensions.AI.ImageGeneratorExtensions.GenerateImagesAsync(Microsoft.Extensions.AI.IImageGenerator,System.String,Microsoft.Extensions.AI.ImageGenerationOptions,System.Threading.CancellationToken)> method with a text prompt. - Saves the generated image to a file in the local user directory. 1. Run the application, either through the IDE or using `dotnet run`. The application generates an image and outputs the file path to the image. Open the file to view the generated image. The following image shows one example of a generated image. ## Configure image generation options You can customize image generation by providing other options such as size, response format, and number of images to generate. The <xref:Microsoft.Extensions.AI.ImageGenerationOptions> class allows you to specify: - <xref:Microsoft.Extensions.AI.ImageGenerationOptions.AdditionalProperties>: Provider-specific options. - <xref:Microsoft.Extensions.AI.ImageGenerationOptions.Count>: The number of images to generate. - <xref:Microsoft.Extensions.AI.ImageGenerationOptions.ImageSize>: The dimensions of the generated image as a <xref:System.Drawing.Size?displayProperty=fullName>. For supported sizes, see the [OpenAI API reference](https://platform.openai.com/docs/api-reference/images/create). - <xref:Microsoft.Extensions.AI.ImageGenerationOptions.MediaType>: The media type (MIME type) of the generated image. - <xref:Microsoft.Extensions.AI.ImageGenerationOptions.ModelId>: The model ID. - <xref:Microsoft.Extensions.AI.ImageGenerationOptions.RawRepresentationFactory>: The callback that creates the raw representation of the image generation options from an underlying implementation. - <xref:Microsoft.Extensions.AI.ImageGenerationOptions.ResponseFormat>: Options are <xref:Microsoft.Extensions.AI.ImageGenerationResponseFormat.Uri>, <xref:Microsoft.Extensions.AI.ImageGenerationResponseFormat.Data>, and <xref:Microsoft.Extensions.AI.ImageGenerationResponseFormat.Hosted>. ## Use hosting integration When you build web apps or hosted services, you can integrate image generation using dependency injection and hosting patterns. This approach provides better lifecycle management, configuration integration, and testability. ### Configure hosting services The `Aspire.Azure.AI.OpenAI` package provides extension methods to register Azure OpenAI services with your application's dependency injection container: 1. Add the necessary packages to your web application: ```dotnetcli dotnet add package Aspire.Azure.AI.OpenAI --prerelease dotnet add package Azure.AI.OpenAI dotnet add package Microsoft.Extensions.AI.OpenAI --prerelease ``` 1. Configure the Azure OpenAI client and image generator in your `Program.cs` file: The <xref:Microsoft.Extensions.Hosting.AspireAzureOpenAIExtensions.AddAzureOpenAIClient(Microsoft.Extensions.Hosting.IHostApplicationBuilder,System.String,System.Action{Aspire.Azure.AI.OpenAI.AzureOpenAISettings},System.Action{Azure.Core.Extensions.IAzureClientBuilder{Azure.AI.OpenAI.AzureOpenAIClient,Azure.AI.OpenAI.AzureOpenAIClientOptions}})> method registers the Azure OpenAI client with dependency injection. The connection string (named `"openai"`) is retrieved from configuration, typically from `appsettings.json` or environment variables: ```json { "ConnectionStrings": { "openai": "Endpoint=https://your-resource-name.openai.azure.com/;Key=your-api-key" } } ``` 1. Register the <xref:Microsoft.Extensions.AI.IImageGenerator> service with dependency injection: The <xref:Microsoft.Extensions.DependencyInjection.ImageGeneratorBuilderServiceCollectionExtensions.AddImageGenerator*> method registers the image generator as a singleton service that can be injected into controllers, services, or minimal API endpoints. 1. Add options and logging:: The preceding code: - Configures options by calling the <xref:Microsoft.Extensions.AI.ConfigureOptionsImageGeneratorBuilderExtensions.ConfigureOptions(Microsoft.Extensions.AI.ImageGeneratorBuilder,System.Action{Microsoft.Extensions.AI.ImageGenerationOptions})> extension method on the <xref:Microsoft.Extensions.AI.ImageGeneratorBuilder>. This method configures the <xref:Microsoft.Extensions.AI.ImageGenerationOptions> to be passed to the next generator in the pipeline. - Adds logging to the image generator pipeline by calling the <xref:Microsoft.Extensions.AI.LoggingImageGeneratorBuilderExtensions.UseLogging(Microsoft.Extensions.AI.ImageGeneratorBuilder,Microsoft.Extensions.Logging.ILoggerFactory,System.Action{Microsoft.Extensions.AI.LoggingImageGenerator})> extension method. ### Use the image generator in endpoints Once registered, you can inject `IImageGenerator` into your endpoints or services: This hosting approach provides several benefits: - **Configuration management**: Connection strings and settings are managed through the .NET configuration system. - **Dependency injection**: The image generator is available throughout your application via DI. - **Lifecycle management**: Services are properly initialized and disposed of by the hosting infrastructure. - **Testability**: Mock implementations can be easily substituted for testing. - **Integration with .NET Aspire**: When using .NET Aspire, the `AddAzureOpenAIClient` method integrates with service discovery and telemetry. ## Best practices When implementing text-to-image generation in your applications, consider these best practices: - **Prompt engineering**: Write clear, detailed prompts that describe the desired image. Include specific details about style, composition, colors, and elements. - **Cost management**: Image generation can be expensive. Cache results when possible and implement rate limiting to control costs. - **Content safety**: Always review generated images for appropriate content, especially in production applications. Consider implementing content filtering and moderation. - **User experience**: Image generation can take several seconds. Provide progress indicators and handle timeouts gracefully. - **Legal considerations**: Be aware of licensing and usage rights for generated images. Review the terms of service for your AI provider. ## Clean up resources When you no longer need the Azure OpenAI resource, delete it to avoid incurring charges: 1. In the [Azure portal](https://portal.azure.com), navigate to your Azure OpenAI resource. 1. Select the resource and then select **Delete**. ## Next steps You've successfully generated some different images using the <xref:Microsoft.Extensions.AI.IImageGenerator> interface in <xref:Microsoft.Extensions.AI>. Next, you can explore some of the additional functionality, including: - Refining the generated image iteratively. - Editing an existing image. - Personalizing an image, diagram, or theme. ## See also - [Explore text-to-image capabilities in .NET (blog post)](https://devblogs.microsoft.com/dotnet/explore-text-to-image-dotnet/) - [Microsoft.Extensions.AI library overview](../microsoft-extensions-ai.md) - [Quickstart: Build an AI chat app with .NET](../quickstarts/build-chat-app.md) -
use-function-calling.md 4.6 KB
--- title: Quickstart - Extend OpenAI using functions and execute a local function with .NET description: Create a simple chat app using OpenAI and extend the model to execute a local function. ms.date: 03/04/2026 ms.topic: quickstart ai-usage: ai-assisted zone_pivot_groups: openai-library --- # Invoke .NET functions using an AI model In this quickstart, you create a .NET console AI chat app that connects to an AI model with local function-calling enabled. The app uses the <xref:Microsoft.Extensions.AI> library so you can write code using AI abstractions rather than a specific SDK. AI abstractions enable you to change the underlying AI model with minimal code changes. :::zone target="docs" pivot="openai" [!INCLUDE [openai-prereqs](includes/prerequisites-openai.md)] :::zone-end :::zone target="docs" pivot="azure-openai" [!INCLUDE [azure-openai-prereqs](includes/prerequisites-azure-openai.md)] :::zone-end ## Create the app Complete the following steps to create a .NET console app to connect to an AI model. 1. In an empty directory on your computer, use the `dotnet new` command to create a new console app: ```dotnetcli dotnet new console -o FunctionCallingAI ``` 1. Change directory into the app folder: ```dotnetcli cd FunctionCallingAI ``` 1. Install the required packages: :::zone target="docs" pivot="azure-openai" ```bash dotnet add package Azure.Identity dotnet add package Azure.AI.OpenAI dotnet add package Microsoft.Extensions.AI dotnet add package Microsoft.Extensions.AI.OpenAI dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` :::zone-end :::zone target="docs" pivot="openai" ```bash dotnet add package Microsoft.Extensions.AI dotnet add package Microsoft.Extensions.AI.OpenAI dotnet add package Microsoft.Extensions.Configuration dotnet add package Microsoft.Extensions.Configuration.UserSecrets ``` :::zone-end 1. Open the app in Visual Studio Code or your editor of choice ```bash code . ``` :::zone target="docs" pivot="azure-openai" [!INCLUDE [create-ai-service](includes/create-ai-service.md)] :::zone-end :::zone target="docs" pivot="openai" ## Configure the app 1. Navigate to the root of your .NET project from a terminal or command prompt. 1. Run the following commands to configure your OpenAI API key as a secret for the sample app: ```bash dotnet user-secrets init dotnet user-secrets set OpenAIKey <your-OpenAI-key> dotnet user-secrets set ModelName <your-OpenAI-model-name> ``` :::zone-end ## Add the app code The app uses the [`Microsoft.Extensions.AI`](https://www.nuget.org/packages/Microsoft.Extensions.AI/) package to send and receive requests to the AI model. 1. In the **Program.cs** file, add the following code to connect and authenticate to the AI model. The `ChatClient` is also configured to use function invocation, which allows the AI model to call .NET functions in your code. :::zone target="docs" pivot="azure-openai" :::zone-end :::zone target="docs" pivot="openai" :::zone-end 1. Create a new `ChatOptions` object that contains an inline function the AI model can call to get the current weather. The function declaration includes a delegate to run logic, and name and description parameters to describe the purpose of the function to the AI model. 1. Add a system prompt to the `chatHistory` to provide context and instructions to the model. Send a user prompt with a question that requires the AI model to call the registered function to properly answer the question. 1. Use the `dotnet run` command to run the app: ```dotnetcli dotnet run ``` The app prints the completion response from the AI model, which includes data provided by the .NET function. The AI model understood that the registered function was available and called it automatically to generate a proper response. :::zone target="docs" pivot="azure-openai" ## Clean up resources If you no longer need them, delete the Azure OpenAI resource and GPT-4 model deployment. 1. In the [Azure portal](https://aka.ms/azureportal), navigate to the Azure OpenAI resource. 1. Select the Azure OpenAI resource, and then select **Delete**. :::zone-end ## Next steps - [Handle invalid tool input from AI models](../how-to/handle-invalid-tool-input.md) - [Access data in AI functions](../how-to/access-data-in-functions.md) - [Quickstart - Build an AI chat app with .NET](build-chat-app.md) - [Generate text and conversations with .NET and Azure OpenAI Completions](/training/modules/open-ai-dotnet-text-completions/)
-
-
resources
-
azure-ai.md 625 B
--- title: Azure AI learning resources description: This article provides a list of resources about Azure AI scenarios for .NET developers, including documentation and code samples. ms.date: 09/04/2025 ms.topic: reference --- # Azure AI learning resources This article contains an organized list of the best learning resources for .NET developers who are building AI apps using Azure services. Resources include popular quickstart articles, reference samples, documentation, and training courses. [!INCLUDE [include-file-from-azure-dev-docs-pr](~/azure-dev-docs-pr/articles/ai/includes/azure-ai-for-developers-dotnet.md)] -
get-started.md 845 B
--- title: Learning resources to get started with AI in .NET description: This article provides a list of resources for .NET developers who are starting to build AI apps with .NET. ms.date: 09/04/2025 ms.topic: reference --- # Learning resources to get started with AI in .NET This article contains an organized list of the best learning resources for .NET developers who are starting to build AI apps with .NET. Resources include samples, documentation, videos, and workshops. ## Tutorials - [Generative AI for beginners](https://github.com/microsoft/Generative-AI-for-beginners-dotnet) ## Workshops - [AI workshop](https://github.com/dotnet-presentations/ai-workshop) - [Steve Sanderson's AI workshop](https://github.com/SteveSandersonMS/dotnet-ai-workshop) ## Sample apps - [AI samples for .NET](https://github.com/dotnet/ai-samples) -
mcp-servers.md 3.7 KB
--- title: MCP server learning resources description: This article provides a list of resources for .NET developers who are building MCP servers. ms.date: 09/04/2025 ms.topic: reference --- # MCP server learning resources This article contains an organized list of the best learning resources for .NET developers who are building Model Context Protocol (MCP) servers. Resources include samples, documentation, videos, and workshops. ## Getting started | Link | Description | |------|-------------| |[Anthropic - Getting Started with MCP](https://modelcontextprotocol.io/docs/getting-started/intro)|Anthropic's official guide to the Model Context Protocol (MCP) and overview on how to get started writing or using an MCP server.| |[Create a minimal MCP server using C# and publish to NuGet](../quickstarts/build-mcp-server.md)|A quick guide to writing an MCP server in VS Code and publishing it to NuGet.| |[MCP for Beginners GitHub repo](https://github.com/microsoft/mcp-for-beginners)|A Microsoft-official GitHub repository with a full, hands-on curriculum for developing MCP servers in .NET, Java, TypeScript, JavaScript, Rust, and Python.| |[Build agents using MCP on Azure](/azure/developer/ai/build-openai-mcp-server-dotnet)| A walkthrough on how to connect an MCP client using Azure OpenAI and MCP server to Azure Container Apps. | ## Libraries | Link | Description | |------------------------------------------------------------------|-------------| | [MCP C# SDK](https://github.com/modelcontextprotocol/csharp-sdk) | Microsoft and Anthropic's official C# SDK for MCP servers and clients. | ## Samples | Link | Description | |------|-------------| |[MCP Workshop GitHub repo](https://github.com/Azure-Samples/mcp-workshop-dotnet)|A sample MCP server and client with corresponding, step-by-step walkthroughs on how to implement them and how MCP works.| ## Documentation | Link | Description | |------|-------------| |[MCP C# SDK Documentation](https://modelcontextprotocol.github.io/csharp-sdk/index.html)|Anthropic's official guide to the Model Context Protocol (MCP) and overview on how to get started writing or using an MCP server.| |[VS Code MCP Developer Guide](https://code.visualstudio.com/api/extension-guides/ai/mcp)| Official documentation for building, debugging, and registering an MCP server in VS Code.| ## Additional resources | Link | Description | |------|-------------| |[MCP Inspector Tool](https://github.com/modelcontextprotocol/inspector)|The Anthropic-official GitHub repo for the MCP Inspector tool, a visual testing and debugging tool for MCP servers.| |[MCP servers for VS Code](https://code.visualstudio.com/mcp)|A hub for quickly installing first-party and third-party MCP servers in VS Code.| |[Third party MCP servers](https://github.com/modelcontextprotocol/servers?tab=readme-ov-file#-third-party-servers)|A GitHub repo containing a comprehensive list of publicly available MCP servers.| ## Videos | Link | Description | |------|-------------| |[MCP Dev Days (Microsoft Reactor)](https://developer.microsoft.com/reactor/series/S-1563/)|A 2-day virtual event and video series showing how to use and write MCP servers.| |[ASP.NET Community Standup - Build MCP servers with ASP.NET Core](https://www.youtube.com/live/x_6iUhdHnhc?si=J8QQuirYWk0JXC_V)|A roundtable discussion and demo on building an MCP server with ASP.NET Core and the MCP C# SDK.| ## Communities | Link | Description | |------|-------------| |[Anthropic MCP Discussions](https://github.com/orgs/modelcontextprotocol/discussions)|An open forum for discussing MCP in Anthropic's official MCP GitHub repo.| |[Microsoft Foundry Discord](https://discord.com/invite/ByRwuEEgH4)|A Discord discussion space for developers looking to build with Microsoft AI tooling.|
-
-
tutorials
-
tutorial-ai-vector-search.md 12.4 KB
--- title: Tutorial - Integrate OpenAI with the RAG pattern and vector search using Azure Cosmos DB for MongoDB description: Create a simple recipe app using the RAG pattern and vector search using Azure Cosmos DB for MongoDB. ms.date: 03/04/2026 ai-usage: ai-assisted ms.topic: tutorial author: alexwolfmsft ms.author: alexwolf --- # Implement Azure OpenAI with RAG using vector search in a .NET app This tutorial explores integration of the RAG pattern using OpenAI models and vector search capabilities in a .NET app. The sample application performs vector searches on custom data stored in Azure Cosmos DB for MongoDB and further refines the responses using generative AI models, such as gpt-5. In the sections that follow, you set up a sample application and explore key code examples that demonstrate these concepts. ## Prerequisites - [.NET 8.0](https://dotnet.microsoft.com/) - An [Azure Account](https://azure.microsoft.com/free) - An [Azure Cosmos DB for MongoDB vCore](/azure/cosmos-db/mongodb/vcore/introduction) service - An [Azure Open AI](/azure/ai-services/openai/overview) service - Deploy `text-embedding-ada-002` model for embeddings - Deploy `gpt-35-turbo` model for chat completions ## App overview The Cosmos Recipe Guide app lets you perform vector and AI-driven searches against a set of recipe data. Search directly for available recipes or prompt the app with ingredient names to find related recipes. The app and the sections ahead guide you through the following workflow to demonstrate this type of functionality: 1. Upload sample data to an Azure Cosmos DB for MongoDB database. 1. Create embeddings and a vector index for the uploaded sample data using the Azure OpenAI `text-embedding-3-small` model. 1. Perform vector similarity search based on the user prompts. 1. Use the Azure OpenAI `gpt-35-turbo` completions model to compose more meaningful answers based on the search results data. ## Get started 1. Clone the following GitHub repository: ```bash git clone https://github.com/microsoft/AzureDataRetrievalAugmentedGenerationSamples.git ``` 1. In the _C#/CosmosDB-MongoDBvCore_ folder, open the **CosmosRecipeGuide.sln** file. 1. In the _appsettings.json_ file, replace the following config values with your Azure OpenAI and Azure Cosmos DB for MongoDB values: ```json "OpenAIEndpoint": "https://<your-service-name>.openai.azure.com/", "OpenAIKey": "<your-API-key>", "OpenAIEmbeddingDeployment": "<your-ADA-deployment-name>", "OpenAIcompletionsDeployment": "<your-GPT-deployment-name>", "MongoVcoreConnection": "<your-Mongo-connection-string>" ``` 1. Launch the app by pressing the **Start** button at the top of Visual Studio. ## Explore the app When you run the app for the first time, it connects to Azure Cosmos DB and reports that there are no recipes available yet. Follow the steps displayed by the app to begin the core workflow. 1. Select **Upload recipe(s) to Cosmos DB** and press <kbd>Enter</kbd>. This command reads sample JSON files from the local project and uploads them to the Cosmos DB account. The code from the _Utility.cs_ class parses the local JSON files. ``` C# public static List<Recipe> ParseDocuments(string Folderpath) { List<Recipe> recipes = new List<Recipe>(); Directory.GetFiles(Folderpath) .ToList() .ForEach(f => { var jsonString= System.IO.File.ReadAllText(f); Recipe recipe = JsonConvert.DeserializeObject<Recipe>(jsonString); recipe.id = recipe.name.ToLower().Replace(" ", ""); recipes.Add(recipe); } ); return recipes; } ``` The `UpsertVectorAsync` method in the _VCoreMongoService.cs_ file uploads the documents to Azure Cosmos DB for MongoDB. ```C# public async Task UpsertVectorAsync(Recipe recipe) { BsonDocument document = recipe.ToBsonDocument(); if (!document.Contains("_id")) { Console.WriteLine("UpsertVectorAsync: Document does not contain _id."); throw new ArgumentException("UpsertVectorAsync: Document does not contain _id."); } string? _idValue = document["_id"].ToString(); try { var filter = Builders<BsonDocument>.Filter.Eq("_id", _idValue); var options = new ReplaceOptions { IsUpsert = true }; await _recipeCollection.ReplaceOneAsync(filter, document, options); } catch (Exception ex) { Console.WriteLine($"Exception: UpsertVectorAsync(): {ex.Message}"); throw; } } ``` 1. Select **Vectorize the recipe(s) and store them in Cosmos DB**. The JSON items uploaded to Cosmos DB don't contain embeddings and therefore are not optimized for RAG via vector search. An embedding is an information-dense, numerical representation of the semantic meaning of a piece of text. Vector searches can find items with contextually similar embeddings. The `GetEmbeddingsAsync` method in the _OpenAIService.cs_ file creates an embedding for each item in the database. ```C# public async Task<float[]?> GetEmbeddingsAsync(dynamic data) { try { EmbeddingsOptions options = new EmbeddingsOptions(data) { Input = data }; var response = await _openAIClient.GetEmbeddingsAsync(openAIEmbeddingDeployment, options); Embeddings embeddings = response.Value; float[] embedding = embeddings.Data[0].Embedding.ToArray(); return embedding; } catch (Exception ex) { Console.WriteLine($"GetEmbeddingsAsync Exception: {ex.Message}"); return null; } } ``` The `CreateVectorIndexIfNotExists` in the _VCoreMongoService.cs_ file creates a vector index, which lets you perform vector similarity searches. ```C# public void CreateVectorIndexIfNotExists(string vectorIndexName) { try { //Find if vector index exists in vectors collection using (IAsyncCursor<BsonDocument> indexCursor = _recipeCollection.Indexes.List()) { bool vectorIndexExists = indexCursor.ToList().Any(x => x["name"] == vectorIndexName); if (!vectorIndexExists) { BsonDocumentCommand<BsonDocument> command = new BsonDocumentCommand<BsonDocument>( BsonDocument.Parse(@" { createIndexes: 'Recipe', indexes: [{ name: 'vectorSearchIndex', key: { embedding: 'cosmosSearch' }, cosmosSearchOptions: { kind: 'vector-ivf', numLists: 5, similarity: 'COS', dimensions: 1536 } }] }")); BsonDocument result = _database.RunCommand(command); if (result["ok"] != 1) { Console.WriteLine("CreateIndex failed with response: " + result.ToJson()); } } } } catch (MongoException ex) { Console.WriteLine("MongoDbService InitializeVectorIndex: " + ex.Message); throw; } } ``` 1. Select the **Ask AI Assistant (search for a recipe by name or description, or ask a question)** option in the app to run a user query. The app converts the user query to an embedding using the OpenAI service and the embedding model, then sends the embedding to Azure Cosmos DB for MongoDB to perform a vector search. The `VectorSearchAsync` method in the _VCoreMongoService.cs_ file performs a vector search to find vectors that are close to the supplied vector and returns a list of documents from Azure Cosmos DB for MongoDB vCore. ```C# public async Task<List<Recipe>> VectorSearchAsync(float[] queryVector) { List<string> retDocs = new List<string>(); string resultDocuments = string.Empty; try { //Search Azure Cosmos DB for MongoDB vCore collection for similar embeddings //Project the fields that are needed BsonDocument[] pipeline = new BsonDocument[] { BsonDocument.Parse( @$"{{$search: {{ cosmosSearch: {{ vector: [{string.Join(',', queryVector)}], path: 'embedding', k: {_maxVectorSearchResults}}}, returnStoredSource:true }} }}"), BsonDocument.Parse($"{{$project: {{embedding: 0}}}}"), }; var bsonDocuments = await _recipeCollection .Aggregate<BsonDocument>(pipeline).ToListAsync(); var recipes = bsonDocuments .ToList() .ConvertAll(bsonDocument => BsonSerializer.Deserialize<Recipe>(bsonDocument)); return recipes; } catch (MongoException ex) { Console.WriteLine($"Exception: VectorSearchAsync(): {ex.Message}"); throw; } } ``` The `GetChatCompletionAsync` method generates an improved chat completion response based on the user prompt and the related vector search results. ``` C# public async Task<(string response, int promptTokens, int responseTokens)> GetChatCompletionAsync(string userPrompt, string documents) { try { ChatMessage systemMessage = new ChatMessage( ChatRole.System, _systemPromptRecipeAssistant + documents); ChatMessage userMessage = new ChatMessage( ChatRole.User, userPrompt); ChatCompletionsOptions options = new() { Messages = { systemMessage, userMessage }, MaxTokens = openAIMaxTokens, Temperature = 0.5f, //0.3f, NucleusSamplingFactor = 0.95f, FrequencyPenalty = 0, PresencePenalty = 0 }; Azure.Response<ChatCompletions> completionsResponse = await openAIClient.GetChatCompletionsAsync(openAICompletionDeployment, options); ChatCompletions completions = completionsResponse.Value; return ( response: completions.Choices[0].Message.Content, promptTokens: completions.Usage.PromptTokens, responseTokens: completions.Usage.CompletionTokens ); } catch (Exception ex) { string message = $"OpenAIService.GetChatCompletionAsync(): {ex.Message}"; Console.WriteLine(message); throw; } } ``` The app also uses prompt engineering to ensure OpenAI service limits and formats the response for supplied recipes. ```C# //System prompts to send with user prompts to instruct the model for chat session private readonly string _systemPromptRecipeAssistant = @" You are an intelligent assistant for Contoso Recipes. You are designed to provide helpful answers to user questions about recipes, cooking instructions provided in JSON format below. Instructions: - Only answer questions related to the recipe provided below. - Don't reference any recipe not provided below. - If you're unsure of an answer, say ""I don't know"" and recommend users search themselves. - Your response should be complete. - List the Name of the Recipe at the start of your response followed by step by step cooking instructions. - Assume the user is not an expert in cooking. - Format the content so that it can be printed to the Command Line console. - In case there is more than one recipe you find, let the user pick the most appropriate recipe."; ```
-
-
azure-ai-services-authentication.md 8.9 KB
--- title: Authenticate to Azure OpenAI using .NET description: Learn about the different options to authenticate to Azure OpenAI and other services using .NET. author: alexwolfmsft ms.topic: concept-article ms.date: 03/06/2026 ai-usage: ai-assisted --- # Foundry tools authentication and authorization using .NET Application requests to Microsoft Foundry tools must be authenticated. In this article, you explore the options available to authenticate to Azure OpenAI and other Foundry tools using .NET. Most Foundry tools offer two primary ways to authenticate apps and users: - **Key-based authentication** provides access to an Azure service using secret key values. These secret values are sometimes known as API keys or access keys depending on the service. - **Microsoft Entra ID** provides a comprehensive identity and access management solution to ensure that the correct identities have the correct level of access to different Azure resources. The sections ahead provide conceptual overviews for these two approaches, rather than detailed implementation steps. For more detailed information about connecting to Azure services, visit the following resources: - [Authenticate .NET apps to Azure services](../azure/sdk/authentication/index.md) - [Identity fundamentals](/entra/fundamentals/identity-fundamental-concepts) - [What is Azure RBAC?](/azure/role-based-access-control/overview) > [!NOTE] > The examples in this article focus primarily on connections to Azure OpenAI, but the same concepts and implementation steps directly apply to many other Foundry tools as well. ## Authentication using keys Access keys allow apps and tools to authenticate to a Foundry tool, such as Azure OpenAI, using a secret key provided by the service. Retrieve the secret key using tools such as the Azure portal or Azure CLI and use it to configure your app code to connect to the Foundry tool: ```csharp builder.Services.AddAzureOpenAIChatCompletion( "deployment-model", "service-endpoint", "service-key"); // Secret key var kernel = builder.Build(); ``` Keys are straightforward to use, but treat them with caution. Keys aren't the recommended authentication option because they: - Don't follow [the principle of least privilege](/entra/identity-platform/secure-least-privileged-access). They provide elevated permissions regardless of who uses them or for what task. - Can accidentally end up in source control or unsafe storage locations. - Can easily be shared with or sent to parties who shouldn't have access. - Often require manual administration and rotation. Instead, consider using [Microsoft Entra ID](#authentication-using-microsoft-entra-id) for authentication, which is the recommended solution for most scenarios. ## Authentication using Microsoft Entra ID Microsoft Entra ID is a cloud-based identity and access management service that provides a vast set of features for different business and app scenarios. Microsoft Entra ID is the recommended solution to connect to Azure OpenAI and other Foundry tools and provides the following benefits: - Keyless authentication using [identities](/entra/fundamentals/identity-fundamental-concepts). - Role-based access control (RBAC) to assign identities the minimum required permissions. - Lets you use the [`Azure.Identity`](/dotnet/api/overview/azure/identity-readme) client library to detect [different credentials across environments](/dotnet/api/azure.identity.defaultazurecredential) without requiring code changes. - Automatically handles administrative maintenance tasks such as rotating underlying keys. The workflow to implement Microsoft Entra authentication in your app generally includes the following steps: - Local development: 1. Sign-in to Azure using a local dev tool such as the Azure CLI or Visual Studio. 1. Configure your code to use the [`Azure.Identity`](/dotnet/api/overview/azure/identity-readme) client library and `DefaultAzureCredential` class. 1. Assign Azure roles to the account you signed-in with to enable access to the Foundry tool. - Azure-hosted app: 1. Deploy the app to Azure after configuring it to authenticate using the `Azure.Identity` client library. 1. Assign a [managed identity](/entra/identity/managed-identities-azure-resources/overview) to the Azure-hosted app. 1. Assign Azure roles to the managed identity to enable access to the Foundry tool. The key concepts of this workflow are explored in the following sections. ### Authenticate to Azure locally When developing apps locally that connect to Foundry tools, authenticate to Azure using a tool such as Visual Studio or the Azure CLI. Your local credentials can be discovered by the `Azure.Identity` client library and used to authenticate your app to Azure services, as described in the [Configure the app code](#configure-the-app-code) section. For example, to authenticate to Azure locally using the Azure CLI, run the following command: ```azurecli az login ``` ### Configure the app code Use the [`Azure.Identity`](/dotnet/api/overview/azure/identity-readme) client library from the Azure SDK to implement Microsoft Entra authentication in your code. The `Azure.Identity` libraries include the `DefaultAzureCredential` class, which automatically discovers available Azure credentials based on the current environment and tooling available. For the full set of supported environment credentials and the order in which `DefaultAzureCredential` searches them, see the [Azure SDK for .NET](/dotnet/api/azure.identity.defaultazurecredential) documentation. For example, configure Azure OpenAI to authenticate using `DefaultAzureCredential` using the following code: ```csharp AzureOpenAIClient azureClient = new( new Uri(endpoint), new DefaultAzureCredential(new DefaultAzureCredentialOptions() { TenantId = tenantId } ) ); ``` `DefaultAzureCredential` enables apps to be promoted from local development to production without code changes. For example, during development `DefaultAzureCredential` uses your local user credentials from Visual Studio or the Azure CLI to authenticate to the Foundry tool. When the app is deployed to Azure, `DefaultAzureCredential` uses the managed identity that is assigned to your app. ### Assign roles to your identity [Azure role-based access control (Azure RBAC)](/azure/role-based-access-control) is a system that provides fine-grained access management of Azure resources. Assign a role to the security principal used by `DefaultAzureCredential` to connect to a Foundry tool, whether that's an individual user, group, service principal, or managed identity. Azure roles are a collection of permissions that allow the identity to perform various tasks, such as generate completions or create and delete resources. Assign roles such as **Cognitive Services OpenAI User** (role ID: `5e0bd9bd-7b93-4f28-af87-19fc36ad61bd`) to the relevant identity using tools such as the Azure CLI, Bicep, or the Azure portal. For example, use the `az role assignment create` command to assign a role using the Azure CLI: ```azurecli az role assignment create \ --role "5e0bd9bd-7b93-4f28-af87-19fc36ad61bd" \ --assignee-object-id "$PRINCIPAL_ID" \ --scope /subscriptions/"$SUBSCRIPTION_ID"/resourceGroups/"$RESOURCE_GROUP" \ --assignee-principal-type User ``` Learn more about Azure RBAC using the following resources: - [What is Azure RBAC?](/azure/role-based-access-control/overview) - [Grant a user access](/azure/role-based-access-control/quickstart-assign-role-user-portal) - [RBAC best practices](/azure/role-based-access-control/best-practices) ### Assign a managed identity to your app In most scenarios, Azure-hosted apps should use a [managed identity](/entra/identity/managed-identities-azure-resources/overview) to connect to other services such as Azure OpenAI. Managed identities provide a fully managed identity in Microsoft Entra ID for apps to use when connecting to resources that support Microsoft Entra authentication. `DefaultAzureCredential` discovers the identity associated with your app and uses it to authenticate to other Azure services. There are two types of managed identities you can assign to your app: - A **system-assigned identity** is tied to your application and is deleted if your app is deleted. An app can only have one system-assigned identity. - A **user-assigned identity** is a standalone Azure resource that can be assigned to your app. An app can have multiple user-assigned identities. Assign roles to a managed identity just like you would an individual user account, such as the **Cognitive Services OpenAI User** role. Learn more about working with managed identities using the following resources: - [Managed identities overview](/entra/identity/managed-identities-azure-resources/overview) - [Authenticate App Service to Azure OpenAI using Microsoft Entra ID](/dotnet/ai/how-to/app-service-aoai-auth?pivots=azure-portal) - [How to use managed identities for App Service and Azure Functions](/azure/app-service/overview-managed-identity) -
dotnet-ai-ecosystem.md 9.6 KB
--- title: .NET + AI ecosystem tools and SDKs description: This article provides an overview of the ecosystem of SDKs and tools available to .NET developers integrating AI into their applications. ms.date: 12/10/2025 ms.topic: overview --- # .NET + AI ecosystem tools and SDKs The .NET ecosystem provides many powerful tools, libraries, and services to develop AI applications. .NET supports both cloud and local AI model connections, many different SDKs for various AI and vector database services, and other tools to help you build intelligent apps of varying scope and complexity. > [!IMPORTANT] > Not all of the SDKs and services presented in this article are maintained by Microsoft. When considering an SDK, make sure to evaluate its quality, licensing, support, and compatibility to ensure they meet your requirements. ## Microsoft.Extensions.AI libraries [`Microsoft.Extensions.AI`](microsoft-extensions-ai.md) is a set of core .NET libraries that provide a unified layer of C# abstractions for interacting with AI services, such as small and large language models (SLMs and LLMs), embeddings, and middleware. These APIs were created in collaboration with developers across the .NET ecosystem. The low-level APIs, such as <xref:Microsoft.Extensions.AI.IChatClient> and <xref:Microsoft.Extensions.AI.IEmbeddingGenerator`2>, were extracted from Semantic Kernel and moved into the <xref:Microsoft.Extensions.AI> namespace. `Microsoft.Extensions.AI` provides abstractions that can be implemented by various services, all adhering to the same core concepts. This library is not intended to provide APIs tailored to any specific provider's services. The goal of `Microsoft.Extensions.AI` is to act as a unifying layer within the .NET ecosystem, enabling developers to choose their preferred frameworks and libraries while ensuring seamless integration and collaboration across the ecosystem. ## Other AI-related Microsoft.Extensions libraries The [📦 Microsoft.Extensions.VectorData.Abstractions package](https://www.nuget.org/packages/Microsoft.Extensions.VectorData.Abstractions/) provides a unified layer of abstractions for interacting with a variety of vector stores. It lets you store processed chunks in vector stores such as Qdrant, Azure SQL, CosmosDB, MongoDB, ElasticSearch, and many more. For more information, see [Build a .NET AI vector search app](quickstarts/build-vector-search-app.md). The [📦 Microsoft.Extensions.DataIngestion package](https://www.nuget.org/packages/Microsoft.Extensions.DataIngestion) provides foundational .NET building blocks for data ingestion. It enables developers to read, process, and prepare documents for AI and machine learning workflows, especially retrieval-augmented generation (RAG) scenarios. For more information, see [Data ingestion](conceptual/data-ingestion.md). ## Microsoft Agent Framework If you want to use low-level services, such as <xref:Microsoft.Extensions.AI.IChatClient> and <xref:Microsoft.Extensions.AI.IEmbeddingGenerator`2>, you can reference the `Microsoft.Extensions.AI.Abstractions` package directly from your app. However, if you want to build agentic AI applications with higher-level orchestration capabilities, you should use [Microsoft Agent Framework](/agent-framework/overview/agent-framework-overview). Agent Framework builds on the `Microsoft.Extensions.AI.Abstractions` package and provides concrete implementations of <xref:Microsoft.Extensions.AI.IChatClient> for different services, including OpenAI, Azure OpenAI, Microsoft Foundry, and more. This framework is the recommended approach for .NET apps that need to build agentic AI systems with advanced orchestration, multi-agent collaboration, and enterprise-grade security and observability. Agent Framework is a production-ready, open-source framework that brings together the best capabilities of Semantic Kernel and Microsoft Research's AutoGen. Agent Framework provides: - **Multi-agent orchestration**: Support for sequential, concurrent, group chat, handoff, and *magentic* (where a lead agent directs other agents) orchestration patterns. - **Cloud and provider flexibility**: Cloud-agnostic (containers, on-premises, or multi-cloud) and provider-agnostic (for example, OpenAI or Foundry) using plugin and connector models. - **Enterprise-grade features**: Built-in observability (OpenTelemetry), Microsoft Entra security integration, and responsible AI features including prompt injection protection and task adherence monitoring. - **Standards-based interoperability**: Integration with open standards like Agent-to-Agent (A2A) protocol and Model Context Protocol (MCP) for agent discovery and tool interaction. For more information, see the [Microsoft Agent Framework documentation](/agent-framework/overview/agent-framework-overview). ## Semantic Kernel for .NET [Semantic Kernel](/semantic-kernel/overview/) is an open-source library that enables AI integration and orchestration capabilities in your .NET apps. However, for new applications that require agentic capabilities, multi-agent orchestration, or enterprise-grade observability and security, the recommended framework is [Microsoft Agent Framework](/agent-framework/overview/agent-framework-overview). ## .NET SDKs for building AI apps Many different SDKs are available to build .NET apps with AI capabilities depending on the target platform or AI model. OpenAI models offer powerful generative AI capabilities, while other Foundry tools provide intelligent solutions for a variety of specific scenarios. ### .NET SDKs for OpenAI models | NuGet package | Supported models | Maintainer or vendor | Documentation | |---------------|------------------|----------------------|--------------| | [Microsoft.Agents.AI.OpenAI](https://www.nuget.org/packages/Microsoft.Agents.AI.OpenAI/) | [OpenAI models](https://platform.openai.com/docs/models/overview)<br/>[Azure OpenAI supported models](/azure/ai-services/openai/concepts/models) | [Microsoft Agent Framework](https://github.com/microsoft/agent-framework) (Microsoft) | [Agent Framework documentation](/agent-framework/overview/agent-framework-overview) | | [Azure OpenAI SDK](https://www.nuget.org/packages/Azure.AI.OpenAI/) | [Azure OpenAI supported models](/azure/ai-services/openai/concepts/models) | [Azure SDK for .NET](https://github.com/Azure/azure-sdk-for-net) (Microsoft) | [Azure OpenAI services documentation](/azure/ai-services/openai/) | | [OpenAI SDK](https://www.nuget.org/packages/OpenAI/) | [OpenAI supported models](https://platform.openai.com/docs/models) | [OpenAI SDK for .NET](https://github.com/openai/openai-dotnet) (OpenAI) | [OpenAI services documentation](https://platform.openai.com/docs/overview) | ### .NET SDKs for Foundry Tools Azure offers many other AI services, such as Foundry Tools, to build specific application capabilities and workflows. Most of these services provide a .NET SDK to integrate their functionality into custom apps. Some of the most commonly used services are shown in the following table. For a complete list of available services and learning resources, see the [Foundry Tools](/azure/ai-services/what-are-ai-services) documentation. | Service | Description | |-----------------------------------|----------------------------------------------| | [Azure AI Search](/azure/search/) | Bring AI-powered cloud search to your mobile and web apps. | | [Content Safety in Foundry Control Plane](/azure/ai-services/content-safety/) | Detect unwanted or offensive content. | | [Azure Document Intelligence in Foundry Tools](/azure/ai-services/document-intelligence/) | Turn documents into intelligent data-driven solutions. | | [Azure Language in Foundry Tools](/azure/ai-services/language-service/) | Build apps with industry-leading natural language understanding capabilities. | | [Azure Speech in Foundry Tools](/azure/ai-services/speech-service/) | Speech to text, text to speech, translation, and speaker recognition. | | [Azure Translator in Foundry Tools](/azure/ai-services/translator/) | AI-powered translation technology with support for more than 100 languages and dialects. | | [Azure Vision in Foundry Tools](/azure/ai-services/computer-vision/) | Analyze content in images and videos. | ## Develop with local AI models .NET apps can also connect to local AI models for many different development scenarios. [Microsoft Agent Framework](https://github.com/microsoft/agent-framework) is the recommended tool to connect to local models using .NET. This framework can connect to many different models hosted across a variety of platforms and abstracts away lower-level implementation details. For example, you can use [Ollama](https://ollama.com/) to [connect to local AI models with .NET](quickstarts/chat-local-model.md), including several small language models (SLMs) developed by Microsoft: | Model | Description | |---------------------|-----------------------------------------------------------| | [phi3 models][phi3] | A family of powerful SLMs with groundbreaking performance at low cost and low latency. | | [orca models][orca] | Research models in tasks such as reasoning over user-provided data, reading comprehension, math problem solving, and text summarization. | > [!NOTE] > The preceding SLMs can also be hosted on other services, such as Azure. ## Next steps - [What is Microsoft Agent Framework?](/agent-framework/overview/agent-framework-overview) - [Quickstart - Summarize text using Azure AI chat app with .NET](quickstarts/prompt-model.md) [phi3]: https://azure.microsoft.com/products/phi-3 [orca]: https://www.microsoft.com/research/project/orca/ -
get-started-app-chat-scaling-with-azure-container-apps.md 2.4 KB
--- title: Scale Azure OpenAI for .NET chat sample using RAG description: Learn how to add load balancing to your application to extend the chat app beyond the Azure OpenAI token and model quota limits. ms.date: 03/04/2026 ai-usage: ai-assisted ms.topic: get-started # CustomerIntent: As a .NET developer new to Azure OpenAI, I want to scale my Azure OpenAI capacity to avoid rate limit errors with Azure Container Apps. --- # Scale Azure OpenAI for .NET chat using RAG with Azure Container Apps [!INCLUDE [aca-load-balancer-intro](~/azure-dev-docs-pr/articles/ai/includes//scaling-load-balancer-introduction-azure-container-apps.md)] ## Prerequisites * Azure subscription. [Create one for free](https://azure.microsoft.com/pricing/purchase-options/azure-account?cid=msft_learn). [Dev containers](https://containers.dev/) are available for both samples, with all dependencies required to complete this article. You can run the dev containers in GitHub Codespaces (in a browser) or locally using Visual Studio Code. #### [Codespaces (recommended)](#tab/github-codespaces) * You only need a [GitHub account](https://www.github.com/login) to use Codespaces. #### [Visual Studio Code](#tab/visual-studio-code) * [Docker Desktop](https://www.docker.com/products/docker-desktop/) - Start Docker Desktop if it's not already running * [Visual Studio Code](https://code.visualstudio.com/) * [Dev Container Extension](https://marketplace.visualstudio.com/items?itemName=ms-vscode-remote.remote-containers) --- [!INCLUDE [scaling-load-balancer-aca-procedure.md](~/azure-dev-docs-pr/articles/ai/includes//scaling-load-balancer-procedure-azure-container-apps.md)] [!INCLUDE [redeployment-procedure](~/azure-dev-docs-pr/articles/ai/includes//redeploy-procedure-chat.md)] [!INCLUDE [logs](~/azure-dev-docs-pr/articles/ai/includes//scaling-load-balancer-logs-azure-container-apps.md)] [!INCLUDE [capacity.md](~/azure-dev-docs-pr/articles/ai/includes//scaling-load-balancer-capacity.md)] [!INCLUDE [aca-cleanup](~/azure-dev-docs-pr/articles/ai/includes//scaling-load-balancer-cleanup-azure-container-apps.md)] ## Sample code Samples used in this article include: * [.NET chat app with RAG](https://github.com/Azure-Samples/azure-search-openai-demo-csharp) * [Load Balancer with Azure Container Apps](https://github.com/Azure-Samples/openai-aca-lb) ## Next step * Use [Azure Load Testing](/azure/load-testing/) to load test your chat app -
get-started-app-chat-template.md 17 KB
--- title: "Get started with the 'chat using your own data sample' for .NET" description: Get started with .NET and search across your own data using a chat app sample implemented using Azure OpenAI Service and Retrieval Augmented Generation (RAG) in Azure AI Search. Easily deploy with Azure Developer CLI. This article uses the Azure AI Reference Template sample. ms.date: 03/04/2026 ms.topic: get-started ai-usage: ai-assisted # CustomerIntent: As a .NET developer new to Azure OpenAI, I want deploy and use sample code to interact with app infused with my own business data so that learn from the sample code. --- # Get started with the 'Chat using your own data sample' for .NET This article shows you how to deploy and run the [Chat with your own data sample for .NET](https://github.com/Azure-Samples/azure-search-openai-demo-csharp). This sample implements a chat app using C#, Azure OpenAI Service, and [Retrieval Augmented Generation (RAG)](/azure/search/retrieval-augmented-generation-overview) in Azure AI Search to get answers about employee benefits at a fictitious company. The employee benefits chat app is seeded with PDF files including an employee handbook, a benefits document and a list of company roles and expectations. * [Demo video](https://aka.ms/azai/net/video) In this article, you: - Deploy a chat app to Azure. - Get answers about employee benefits. - Change settings to change behavior of responses. Once you complete this procedure, start modifying the new project with your custom code. This article is part of a collection of articles that show you how to build a chat app using Azure OpenAI service and Azure AI Search. Other articles in the collection include: - [Python](/azure/developer/python/get-started-app-chat-template) - [JavaScript](/azure/developer/javascript/get-started-app-chat-template) - [Java](/azure/developer/java/quickstarts/get-started-app-chat-template) ## Architectural overview In this sample application, a fictitious company called Contoso Electronics provides the chat app experience to its employees to ask questions about the benefits, internal policies, and job descriptions and roles. The architecture of the chat app is shown in the following diagram: - **User interface** - The application's chat interface is a [Blazor WebAssembly](/aspnet/core/blazor/) application. This interface is what accepts user queries, routes request to the application backend, and displays generated responses. - **Backend** - The application backend is an [ASP.NET Core Minimal API](/aspnet/core/fundamentals/minimal-apis/overview). The backend hosts the Blazor static web application and is what orchestrates the interactions among the different services. Services used in this application include: - [**Azure AI Search**](/azure/search/search-what-is-azure-search) – Indexes documents from the data stored in an Azure Storage Account. This makes the documents searchable using [vector search](/azure/search/search-get-started-vector) capabilities. - [**Azure OpenAI Service**](/azure/ai-services/openai/overview) – Provides the Large Language Models (LLM) to generate responses. [Microsoft Agent Framework](/agent-framework/overview/agent-framework-overview) is used in conjunction with the Azure OpenAI Service to orchestrate the more complex AI workflows. ## Cost Most resources in this architecture use a basic or consumption pricing tier. Consumption pricing is based on usage, which means you only pay for what you use. To complete this article, there's a charge, but it's minimal. When you're done with the article, delete the resources to stop incurring charges. For more information, see [Azure Samples: Cost in the sample repo](https://github.com/Azure-Samples/azure-search-openai-demo-csharp#cost-estimation). ## Prerequisites A [development container](https://containers.dev/) environment is available with all dependencies required to complete this article. You can run the development container in GitHub Codespaces (in a browser) or locally using Visual Studio Code. To follow along with this article, you need the following prerequisites: #### [Codespaces (recommended)](#tab/github-codespaces) * An Azure subscription - [Create one for free](https://azure.microsoft.com/pricing/purchase-options/azure-account?cid=msft_learn) * Azure account permissions - Your Azure account must have Microsoft.Authorization/roleAssignments/write permissions, such as [User Access Administrator](/azure/role-based-access-control/built-in-roles#user-access-administrator) or [Owner](/azure/role-based-access-control/built-in-roles#owner). * GitHub account #### [Visual Studio Code](#tab/visual-studio-code) * An Azure subscription - [Create one for free](https://azure.microsoft.com/pricing/purchase-options/azure-account?cid=msft_learn) * Azure account permissions - Your Azure account must have Microsoft.Authorization/roleAssignments/write permissions, such as [User Access Administrator](/azure/role-based-access-control/built-in-roles#user-access-administrator) or [Owner](/azure/role-based-access-control/built-in-roles#owner). * [Azure Developer CLI](/azure/developer/azure-developer-cli) * [Docker Desktop](https://www.docker.com/products/docker-desktop/) - Start Docker Desktop if it's not already running * [Visual Studio Code](https://code.visualstudio.com/) * [Dev Container Extension](https://marketplace.visualstudio.com/items?itemName=ms-vscode-remote.remote-containers) --- ## Open development environment Begin now with a development environment that has all the dependencies installed to complete this article. #### [GitHub Codespaces (recommended)](#tab/github-codespaces) [GitHub Codespaces](https://docs.github.com/codespaces) runs a development container managed by GitHub with [Visual Studio Code for the Web](https://code.visualstudio.com/docs/editor/vscode-web) as the user interface. For the most straightforward development environment, use GitHub Codespaces so that you have the correct developer tools and dependencies preinstalled to complete this article. > [!IMPORTANT] > All GitHub accounts can use Codespaces for up to 60 hours free each month with 2 core instances. For more information, see [GitHub Codespaces monthly included storage and core hours](https://docs.github.com/billing/managing-billing-for-github-codespaces/about-billing-for-github-codespaces#monthly-included-storage-and-core-hours-for-personal-accounts). 1. Start the process to create a new GitHub codespace on the `main` branch of the [`Azure-Samples/azure-search-openai-demo-csharp`](https://github.com/Azure-Samples/azure-search-openai-demo-csharp) GitHub repository. 1. To have both the development environment and the documentation available at the same time, right-click on the following **Open in GitHub Codespaces** button, and select _Open link in new windows_. [](https://codespaces.new/Azure-Samples/azure-search-openai-demo-csharp) 1. On the **Create codespace** page, review the codespace configuration settings and then select **Create new codespace**: 1. Wait for the codespace to start. This startup process can take a few minutes. 1. In the terminal at the bottom of the screen, sign in to Azure with the Azure Developer CLI. ```bash azd auth login ``` 1. Copy the code from the terminal and then paste it into a browser. Follow the instructions to authenticate with your Azure account. 1. The remaining tasks in this article take place in the context of this development container. #### [Visual Studio Code](#tab/visual-studio-code) The [Dev Containers extension](https://marketplace.visualstudio.com/items?itemName=ms-vscode-remote.remote-containers) for Visual Studio Code requires [Docker](https://docs.docker.com/) to be installed on your local machine. The extension hosts the development container locally using the Docker host with the correct developer tools and dependencies preinstalled to complete this article. 1. Create a new local directory on your computer for the project. ```bash mkdir my-intelligent-app && cd my-intelligent-app ``` 1. Open Visual Studio Code in that directory: ```bash code . ``` 1. Open a new terminal in Visual Studio Code. 1. Run the following `azd` command to clone the GitHub repository to your local computer. ```bash azd init -t azure-search-openai-demo-csharp ``` 1. Open the Command Palette, and then search for and select **Dev Containers: Open Folder in Container** to open the project in a dev container. Wait until the dev container opens before continuing. 1. Sign in to Azure with the Azure Developer CLI. ```bash azd auth login ``` Copy the code from the terminal and then paste it into a browser. Follow the instructions to authenticate with your Azure account. 1. The remaining exercises in this project take place in the context of this development container. --- ## Deploy and run The sample repository contains all the code and configuration files you need to deploy a chat app to Azure. The following steps walk you through the process of deploying the sample to Azure. ### Deploy chat app to Azure > [!IMPORTANT] > Azure resources created in this section incur immediate costs, primarily from the Azure AI Search resource. These resources may accrue costs even if you interrupt the command before it's fully executed. 1. Run the following Azure Developer CLI command to provision the Azure resources and deploy the source code: ```bash azd up ``` 1. When you're prompted to enter an environment name, keep it short and lowercase. For example, `myenv`. It's used as part of the resource group name. 1. When prompted, select a subscription to create the resources in. 1. When you're prompted to select a location the first time, select a location near you. This location is used for most the resources including hosting. 1. If you're prompted for a location for the OpenAI model, select a location that is near you. If the same location is available as your first location, select that. 1. Wait until the app is deployed. The deployment might take up to 20 minutes to complete. 1. After the application deploys successfully, a URL appears in the terminal. 1. Select that URL labeled `Deploying service web` to open the chat application in a browser. ### Use chat app to get answers from PDF files The chat app is preloaded with employee benefits information from [PDF files](https://github.com/Azure-Samples/azure-search-openai-demo-csharp/tree/main/data). You can use the chat app to ask questions about the benefits. The following steps walk you through the process of using the chat app. 1. In the browser, navigate to the **Chat** page using the left navigation. 1. Select or enter "What is included in my Northwind Health Plus plan that is not in standard?" in the chat text box. Your response is _similar_ to the following image. 1. From the answer, select a citation. A pop-up window opens displaying the source of the information. 1. Navigate between the tabs at the top of the answer box to understand how the answer was generated. | Tab | Description | |------------------------|-------------| | **Thought process** | This is a script of the interactions in chat. You can view the system prompt (`content`) and your user question (`content`). | | **Supporting content** | This includes the information to answer your question and the source material. The number of source material citations is noted in the **Developer settings**. The default value is **3**. | | **Citation** | This displays the source page that contains the citation. | 1. When you're done, navigate back to the answer tab. ### Use chat app settings to change behavior of responses The intelligence of the chat is determined by the OpenAI model and the settings that are used to interact with the model. | Setting | Description | |-----------------------------|--------------------------------------------------------------------| | Override prompt template | This is the prompt that is used to generate the answer. | | Retrieve this many search results |This is the number of search results that are used to generate the answer. You can see these sources returned in the _Thought process_ and _Supporting content_ tabs of the citation. | | Exclude category | This is the category of documents that are excluded from the search results. | | Use semantic ranker for retrieval | This is a feature of [Azure AI Search](/azure/search/semantic-search-overview#what-is-semantic-search) that uses machine learning to improve the relevance of search results. | | Retrieval mode | **Vectors + Text** means that the search results are based on the text of the documents and the embeddings of the documents. **Vectors** means that the search results are based on the embeddings of the documents. **Text** means that the search results are based on the text of the documents. | | Use query-contextual summaries instead of whole documents | When both `Use semantic ranker` and `Use query-contextual summaries` are checked, the LLM uses captions extracted from key passages, instead of all the passages, in the highest ranked documents. | | Suggest follow-up questions | Have the chat app suggest follow-up questions based on the answer. | The following steps walk you through the process of changing the settings. 1. In the browser, select the gear icon in the upper right of the page. 1. If not selected, select the **Suggest follow-up questions** checkbox and ask the same question again. ```Text What is included in my Northwind Health Plus plan that is not in standard? ``` The chat might return with follow-up question suggestions. 1. In the **Settings** tab, deselect **Use semantic ranker for retrieval**. 1. Ask the same question again. ```Text What is my deductible? ``` 1. What is the difference in the answers? The response that used the Semantic ranker provided a single answer. The response without semantic ranking returned a less direct answer. ## Clean up resources To finish, clean up the Azure and GitHub CodeSpaces resources you used. ### Clean up Azure resources The Azure resources created in this article are billed to your Azure subscription. If you don't expect to need these resources in the future, delete them to avoid incurring more charges. Run the following Azure Developer CLI command to delete the Azure resources and remove the source code: ```bash azd down --purge ``` ### Clean up GitHub Codespaces #### [GitHub Codespaces](#tab/github-codespaces) Deleting the GitHub Codespaces environment ensures that you can maximize the amount of free per-core hours entitlement you get for your account. > [!IMPORTANT] > For more information about your GitHub account's entitlements, see [GitHub Codespaces monthly included storage and core hours](https://docs.github.com/billing/managing-billing-for-github-codespaces/about-billing-for-github-codespaces#monthly-included-storage-and-core-hours-for-personal-accounts). 1. Sign into the GitHub Codespaces dashboard (<https://github.com/codespaces>). 1. Locate your currently running codespaces sourced from the [`Azure-Samples/azure-search-openai-demo-csharp`](https://github.com/Azure-Samples/azure-search-openai-demo-csharp) GitHub repository. 1. Open the context menu for the codespace and then select **Delete**. #### [Visual Studio Code](#tab/visual-studio-code) You don't need to clean up your local environment, but you can stop the running development container and return to running Visual Studio Code locally. 1. Open the **Command Palette**, search for the **Dev Containers** commands, and then select **Dev Containers: Reopen Folder Locally**. > [!TIP] > Visual Studio Code stops the running development container, but the container still exists in Docker in a stopped state. You can also delete the container instance, container image, and volumes from Docker to free up more space on your local machine. --- ## Get help This sample repository offers [troubleshooting information](https://github.com/Azure-Samples/azure-search-openai-demo-csharp/tree/main#troubleshooting). If your issue isn't addressed, log your issue to the repository's [Issues](https://github.com/Azure-Samples/azure-search-openai-demo-csharp/issues). ## Next steps - [Get the source code for the sample used in this article](https://github.com/Azure-Samples/azure-search-openai-demo-csharp) - [Build a chat app with Azure OpenAI](https://aka.ms/azai/chat) best practice solution architecture - [Access control in Generative AI Apps with Azure AI Search](https://techcommunity.microsoft.com/t5/azure-ai-services-blog/access-control-in-generative-ai-applications-with-azure/ba-p/3956408) - [Build an Enterprise ready OpenAI solution with Azure API Management](https://techcommunity.microsoft.com/t5/apps-on-azure-blog/build-an-enterprise-ready-azure-openai-solution-with-azure-api/bc-p/3935407) - [Outperforming vector search with hybrid retrieval and ranking capabilities](https://techcommunity.microsoft.com/t5/azure-ai-services-blog/azure-cognitive-search-outperforming-vector-search-with-hybrid/ba-p/3929167) -
get-started-mcp.md 7 KB
--- title: Get started with .NET AI and MCP description: Learn about .NET AI and MCP key concepts and development resources to get started building MCP clients and servers ms.date: 11/20/2025 ms.topic: overview author: alexwolfmsft ms.author: alexwolf --- # Get started with .NET AI and the Model Context Protocol Model Context Protocol (MCP) is an open protocol designed to standardize integrations between AI apps and external tools and data sources. By using MCP, developers can enhance the capabilities of AI models, enabling them to produce more accurate, relevant, and context-aware responses. For example, using MCP, you can connect your LLM to resources such as: - Document databases or storage services. - Web APIs that expose business data or logic. - Tools that manage files or performing local tasks on a user's device. Many Microsoft products already support MCP, including: - [Copilot Studio](https://www.microsoft.com/microsoft-copilot/blog/copilot-studio/introducing-model-context-protocol-mcp-in-copilot-studio-simplified-integration-with-ai-apps-and-agents/) - [Visual Studio Code GitHub Copilot agent mode](https://code.visualstudio.com/blogs/2025/02/24/introducing-copilot-agent-mode) - [Agent Framework](/agent-framework/user-guide/model-context-protocol/using-mcp-tools) You can use the [MCP C# SDK](#develop-with-the-mcp-c-sdk) to quickly create your own MCP integrations and switch between different AI models without significant code changes. ## MCP client-server architecture MCP uses a client-server architecture that enables an AI-powered app (the host) to connect to multiple MCP servers through MCP clients: - **MCP hosts**: AI tools, code editors, or other software that enhance their AI models using contextual resources through MCP. For example, GitHub Copilot in Visual Studio Code can act as an MCP host and use MCP clients and servers to expand its capabilities. - **MCP clients**: Clients used by the host application to connect to MCP servers to retrieve contextual data. - **MCP servers**: Services that expose capabilities to clients through MCP. For example, an MCP server might provide an abstraction over a REST API or local data source to provide business data to the AI model. The following diagram illustrates this architecture: MCP client and server can exchange a set of standard messages: | Message | Description | |---------------------|---------------------------------------------------------------| | `InitializeRequest` | This request is sent by the client to the server when it first connects, asking it to begin initialization. | | `ListToolsRequest` | Sent by the client to request a list of tools the server has. | | `CallToolRequest` | Used by the client to invoke a tool provided by the server. | | `ListResourcesRequest` | Sent by the client to request a list of available server resources. | | `ReadResourceRequest` | Sent by the client to the server to read a specific resource URI. | | `ListPromptsRequest` | Sent by the client to request a list of available prompts and prompt templates from the server. | | `GetPromptRequest` | Used by the client to get a prompt provided by the server. | | `PingRequest` | A ping, issued by either the server or the client, to check that the other party is still alive. | | `CreateMessageRequest` | A request by the server to sample an LLM via the client. The client has full discretion over which model to select. The client should also inform the user before beginning sampling, to allow them to inspect the request (human in the loop) and decide whether to approve it. | | `SetLevelRequest` | A request by the client to the server, to enable or adjust logging. | ## Develop with the MCP C# SDK As a .NET developer, you can use MCP by creating MCP clients and servers to enhance your apps with custom integrations. MCP reduces the complexity involved in connecting an AI model to various tools, services, and data sources. The official [MCP C# SDK](https://github.com/modelcontextprotocol/csharp-sdk) is available through NuGet and enables you to build MCP clients and servers for .NET apps and libraries. The SDK is maintained through collaboration between Microsoft, Anthropic, and the MCP open protocol organization. To get started, add the MCP C# SDK to your project: ```dotnetcli dotnet add package ModelContextProtocol --prerelease ``` Instead of building unique connectors for each integration point, you can often leverage or reference prebuilt integrations from various providers such as GitHub and Docker: - [Available MCP clients](https://modelcontextprotocol.io/clients) - [Available MCP servers](https://modelcontextprotocol.io/examples) ### Integration with Microsoft.Extensions.AI The MCP C# SDK depends on the [Microsoft.Extensions.AI libraries](/dotnet/ai/ai-extensions) to handle various AI interactions and tasks. These extension libraries provides core types and abstractions for working with AI services, so developers can focus on coding against conceptual AI capabilities rather than specific platforms or provider implementations. View the MCP C# SDK dependencies on the [NuGet package page](https://www.nuget.org/packages/ModelContextProtocol). ## More .NET MCP development resources Various tools, services, and learning resources are available in the .NET and Azure ecosystems to help you build MCP clients and servers or integrate with existing MCP servers. Get started with the following development tools: - [Agent Framework](/agent-framework/user-guide/model-context-protocol/using-mcp-tools) supports integration with MCP servers, allowing your agents to access external tools and services. Agent Framework works with the official MCP C# SDK to enable agents to connect to MCP servers, retrieve available tools, and use them through function calling to extend agent capabilities with external data sources and services. - [Azure Functions remote MCP servers](https://devblogs.microsoft.com/dotnet/build-mcp-remote-servers-with-azure-functions/) combine MCP standards with the flexible architecture of Azure Functions. Visit the [Remote MCP functions sample repository](https://aka.ms/cadotnet/mcp/functions/remote-sample) for code examples. - [Azure MCP Server](https://github.com/Azure/azure-mcp) implements the MCP specification to seamlessly connect AI agents with key Azure services like Azure Storage, Cosmos DB, and more. ## See also - [MCP C# SDK documentation](https://modelcontextprotocol.github.io/csharp-sdk/index.html) - [MCP C# SDK API documentation](https://modelcontextprotocol.github.io/csharp-sdk/api/ModelContextProtocol.html) - [MCP C# SDK README](https://github.com/modelcontextprotocol/csharp-sdk/blob/main/README.md) - [Microsoft partners with Anthropic to create official C# SDK for Model Context Protocol](https://devblogs.microsoft.com/blog/microsoft-partners-with-anthropic-to-create-official-c-sdk-for-model-context-protocol) - [Build a Model Context Protocol (MCP) server in C#](https://devblogs.microsoft.com/dotnet/build-a-model-context-protocol-mcp-server-in-csharp/) -
ichatclient.md 16.5 KB
--- title: Use the IChatClient interface description: Learn how to use the IChatClient interface to get model responses and call tools. ms.date: 03/13/2026 no-loc: ["IChatClient"] --- # Use the IChatClient interface The <xref:Microsoft.Extensions.AI.IChatClient> interface defines a client abstraction responsible for interacting with AI services that provide chat capabilities. It includes methods for sending and receiving messages with multi-modal content (such as text, images, and audio), either as a complete set or streamed incrementally. Additionally, it allows for retrieving strongly typed services provided by the client or its underlying services. .NET libraries that provide clients for language models and services can provide an implementation of the `IChatClient` interface. Any consumers of the interface are then able to interoperate seamlessly with these models and services via the abstractions. You can find examples in the [Implementation examples](#implementation-examples) section. ## Request a chat response With an instance of <xref:Microsoft.Extensions.AI.IChatClient>, you can call the <xref:Microsoft.Extensions.AI.IChatClient.GetResponseAsync*?displayProperty=nameWithType> method to send a request and get a response. The request is composed of one or more messages, each of which is composed of one or more pieces of content. Accelerator methods exist to simplify common cases, such as constructing a request for a single piece of text content. The core `IChatClient.GetResponseAsync` method accepts a list of messages. This list represents the history of all messages that are part of the conversation. The <xref:Microsoft.Extensions.AI.ChatResponse> that's returned from `GetResponseAsync` exposes a list of <xref:Microsoft.Extensions.AI.ChatMessage> instances that represent one or more messages generated as part of the operation. In common cases, there is only one response message, but in some situations, there can be multiple messages. The message list is ordered, such that the last message in the list represents the final message to the request. To provide all of those response messages back to the service in a subsequent request, you can add the messages from the response back into the messages list. ## Request a streaming chat response The inputs to <xref:Microsoft.Extensions.AI.IChatClient.GetStreamingResponseAsync*?displayProperty=nameWithType> are identical to those of `GetResponseAsync`. However, rather than returning the complete response as part of a <xref:Microsoft.Extensions.AI.ChatResponse> object, the method returns an <xref:System.Collections.Generic.IAsyncEnumerable`1> where `T` is <xref:Microsoft.Extensions.AI.ChatResponseUpdate>, providing a stream of updates that collectively form the single response. > [!TIP] > Streaming APIs are nearly synonymous with AI user experiences. C# enables compelling scenarios with its `IAsyncEnumerable<T>` support, allowing for a natural and efficient way to stream data. As with `GetResponseAsync`, you can add the updates from <xref:Microsoft.Extensions.AI.IChatClient.GetStreamingResponseAsync*?displayProperty=nameWithType> back into the messages list. Because the updates are individual pieces of a response, you can use helpers like <xref:Microsoft.Extensions.AI.ChatResponseExtensions.ToChatResponse(System.Collections.Generic.IEnumerable{Microsoft.Extensions.AI.ChatResponseUpdate})> to compose one or more updates back into a single <xref:Microsoft.Extensions.AI.ChatResponse> instance. Helpers like <xref:Microsoft.Extensions.AI.ChatResponseExtensions.AddMessages*> compose a <xref:Microsoft.Extensions.AI.ChatResponse> and then extract the composed messages from the response and add them to a list. ## Tool calling Some models and services support _tool calling_. To gather additional information, you can configure the <xref:Microsoft.Extensions.AI.ChatOptions> with information about tools (usually .NET methods) that the model can request the client to invoke. Instead of sending a final response, the model requests a function invocation with specific arguments. The client then invokes the function and sends the results back to the model with the conversation history. The `Microsoft.Extensions.AI.Abstractions` library includes abstractions for various message content types, including function call requests and results. While `IChatClient` consumers can interact with this content directly, `Microsoft.Extensions.AI` provides helpers that can enable automatically invoking the tools in response to corresponding requests. The `Microsoft.Extensions.AI.Abstractions` and `Microsoft.Extensions.AI` libraries provide the following types: - <xref:Microsoft.Extensions.AI.AIFunction>: Represents a function that can be described to an AI model and invoked. - <xref:Microsoft.Extensions.AI.AIFunctionFactory>: Provides factory methods for creating `AIFunction` instances that represent .NET methods. - <xref:Microsoft.Extensions.AI.FunctionInvokingChatClient>: Wraps an `IChatClient` as another `IChatClient` that adds automatic function-invocation capabilities. The following example demonstrates a random function invocation (this example depends on the [📦 OllamaSharp](https://www.nuget.org/packages/OllamaSharp) NuGet package): The preceding code: - Defines a function named `GetCurrentWeather` that returns a random weather forecast. - Instantiates a <xref:Microsoft.Extensions.AI.ChatClientBuilder> with an `OllamaSharp.OllamaApiClient` and configures it to use function invocation. - Calls `GetStreamingResponseAsync` on the client, passing a prompt and a list of tools that includes a function created with <xref:Microsoft.Extensions.AI.AIFunctionFactory.Create*>. - Iterates over the response, printing each update to the console. For more information about creating AI functions, see [Access data in AI functions](how-to/access-data-in-functions.md). You can also use Model Context Protocol (MCP) tools with your `IChatClient`. For more information, see [Build a minimal MCP client](./quickstarts/build-mcp-client.md). ## Cache responses If you're familiar with [caching in .NET](../core/extensions/caching.md), it's good to know that <xref:Microsoft.Extensions.AI> provides delegating `IChatClient` implementations for caching. The <xref:Microsoft.Extensions.AI.DistributedCachingChatClient> is an `IChatClient` that layers caching around another arbitrary `IChatClient` instance. When a novel chat history is submitted to the `DistributedCachingChatClient`, it forwards it to the underlying client and then caches the response before sending it back to the consumer. The next time the same history is submitted, such that a cached response can be found in the cache, the `DistributedCachingChatClient` returns the cached response rather than forwarding the request along the pipeline. This example depends on the [📦 Microsoft.Extensions.Caching.Memory](https://www.nuget.org/packages/Microsoft.Extensions.Caching.Memory) NuGet package. For more information, see [Caching in .NET](../core/extensions/caching.md). ## Use telemetry Another example of a delegating chat client is the <xref:Microsoft.Extensions.AI.OpenTelemetryChatClient>. This implementation adheres to the [OpenTelemetry Semantic Conventions for Generative AI systems](https://opentelemetry.io/docs/specs/semconv/gen-ai/). Similar to other `IChatClient` delegators, it layers metrics and spans around other arbitrary `IChatClient` implementations. (The preceding example depends on the [📦 OpenTelemetry.Exporter.Console](https://www.nuget.org/packages/OpenTelemetry.Exporter.Console) NuGet package.) Alternatively, the <xref:Microsoft.Extensions.AI.LoggingChatClient> and corresponding <xref:Microsoft.Extensions.AI.LoggingChatClientBuilderExtensions.UseLogging(Microsoft.Extensions.AI.ChatClientBuilder,Microsoft.Extensions.Logging.ILoggerFactory,System.Action{Microsoft.Extensions.AI.LoggingChatClient})> method provide a simple way to write log entries to an <xref:Microsoft.Extensions.Logging.ILogger> for every request and response. ## Provide options Every call to <xref:Microsoft.Extensions.AI.IChatClient.GetResponseAsync*> or <xref:Microsoft.Extensions.AI.IChatClient.GetStreamingResponseAsync*> can optionally supply a <xref:Microsoft.Extensions.AI.ChatOptions> instance containing additional parameters for the operation. The most common parameters among AI models and services show up as strongly typed properties on the type, such as <xref:Microsoft.Extensions.AI.ChatOptions.Temperature?displayProperty=nameWithType>. Other parameters can be supplied by name in a weakly typed manner, via the <xref:Microsoft.Extensions.AI.ChatOptions.AdditionalProperties?displayProperty=nameWithType> dictionary, or via an options instance that the underlying provider understands, using the <xref:Microsoft.Extensions.AI.ChatOptions.RawRepresentationFactory?displayProperty=nameWithType> property. You can also specify options when building an `IChatClient` with the fluent <xref:Microsoft.Extensions.AI.ChatClientBuilder> API by chaining a call to the <xref:Microsoft.Extensions.AI.ConfigureOptionsChatClientBuilderExtensions.ConfigureOptions(Microsoft.Extensions.AI.ChatClientBuilder,System.Action{Microsoft.Extensions.AI.ChatOptions})> extension method. This delegating client wraps another client and invokes the supplied delegate to populate a `ChatOptions` instance for every call. For example, to ensure that the <xref:Microsoft.Extensions.AI.ChatOptions.ModelId?displayProperty=nameWithType> property defaults to a particular model name, you can use code like the following: ## Functionality pipelines `IChatClient` instances can be layered to create a pipeline of components that each add additional functionality. These components can come from `Microsoft.Extensions.AI`, other NuGet packages, or custom implementations. This approach allows you to augment the behavior of the `IChatClient` in various ways to meet your specific needs. Consider the following code snippet that layers a distributed cache, function invocation, and OpenTelemetry tracing around a sample chat client: ## Custom `IChatClient` middleware To add additional functionality, you can implement `IChatClient` directly or use the <xref:Microsoft.Extensions.AI.DelegatingChatClient> class. This class serves as a base for creating chat clients that delegate operations to another `IChatClient` instance. It simplifies chaining multiple clients, which allows calls to pass through to an underlying client. The `DelegatingChatClient` class provides default implementations for methods like `GetResponseAsync`, `GetStreamingResponseAsync`, and `Dispose`, which forward calls to the inner client. A derived class can then override only the methods it needs to augment the behavior, while delegating other calls to the base implementation. This approach is useful for creating flexible and modular chat clients that are easy to extend and compose. The following is an example class derived from `DelegatingChatClient` that uses the [System.Threading.RateLimiting](https://www.nuget.org/packages/System.Threading.RateLimiting) library to provide rate-limiting functionality. As with other `IChatClient` implementations, the `RateLimitingChatClient` can be composed: To simplify the composition of such components with others, component authors should create a `Use*` extension method for registering the component into a pipeline. For example, consider the following `UseRateLimiting` extension method: Such extensions can also query for relevant services from the DI container; the <xref:System.IServiceProvider> used by the pipeline is passed in as an optional parameter: Now it's easy for the consumer to use this in their pipeline, for example: The previous extension methods demonstrate using a `Use` method on <xref:Microsoft.Extensions.AI.ChatClientBuilder>. `ChatClientBuilder` also provides <xref:Microsoft.Extensions.AI.ChatClientBuilder.Use*> overloads that make it easier to write such delegating handlers. For example, in the earlier `RateLimitingChatClient` example, the overrides of `GetResponseAsync` and `GetStreamingResponseAsync` only need to do work before and after delegating to the next client in the pipeline. To achieve the same thing without writing a custom class, you can use an overload of `Use` that accepts a delegate that's used for both `GetResponseAsync` and `GetStreamingResponseAsync`, reducing the boilerplate required: For scenarios where you need a different implementation for `GetResponseAsync` and `GetStreamingResponseAsync` to handle their unique return types, you can use the <xref:Microsoft.Extensions.AI.ChatClientBuilder.Use(System.Func{System.Collections.Generic.IEnumerable{Microsoft.Extensions.AI.ChatMessage},Microsoft.Extensions.AI.ChatOptions,Microsoft.Extensions.AI.IChatClient,System.Threading.CancellationToken,System.Threading.Tasks.Task{Microsoft.Extensions.AI.ChatResponse}},System.Func{System.Collections.Generic.IEnumerable{Microsoft.Extensions.AI.ChatMessage},Microsoft.Extensions.AI.ChatOptions,Microsoft.Extensions.AI.IChatClient,System.Threading.CancellationToken,System.Collections.Generic.IAsyncEnumerable{Microsoft.Extensions.AI.ChatResponseUpdate}})> overload that accepts a delegate for each. ## Dependency injection <xref:Microsoft.Extensions.AI.IChatClient> implementations are often provided to an application via [dependency injection (DI)](../core/extensions/dependency-injection/overview.md). In the following example, an <xref:Microsoft.Extensions.Caching.Distributed.IDistributedCache> is added into the DI container, as is an `IChatClient`. The registration for the `IChatClient` uses a builder that creates a pipeline containing a caching client (which then uses an `IDistributedCache` retrieved from DI) and the sample client. The injected `IChatClient` can be retrieved and used elsewhere in the app. What instance and configuration is injected can differ based on the current needs of the application, and multiple pipelines can be injected with different keys. ## Stateless vs. stateful clients _Stateless_ services require all relevant conversation history to be sent back on every request. In contrast, _stateful_ services keep track of the history and require only additional messages to be sent with a request. The <xref:Microsoft.Extensions.AI.IChatClient> interface is designed to handle both stateless and stateful AI services. When working with a stateless service, callers maintain a list of all messages. They add in all received response messages and provide the list back on subsequent interactions. For stateful services, you might already know the identifier used for the relevant conversation. You can put that identifier into <xref:Microsoft.Extensions.AI.ChatOptions.ConversationId?displayProperty=nameWithType>. Usage then follows the same pattern, except there's no need to maintain a history manually. Some services might support automatically creating a conversation ID for a request that doesn't have one, or creating a new conversation ID that represents the current state of the conversation after incorporating the last round of messages. In such cases, you can transfer the <xref:Microsoft.Extensions.AI.ChatResponse.ConversationId?displayProperty=nameWithType> over to the `ChatOptions.ConversationId` for subsequent requests. For example: If you don't know ahead of time whether the service is stateless or stateful, you can check the response <xref:Microsoft.Extensions.AI.ChatResponse.ConversationId> and act based on its value. If it's set, then that value is propagated to the options and the history is cleared so as to not resend the same history again. If the response `ConversationId` isn't set, then the response message is added to the history so that it's sent back to the service on the next turn. ## Implementation examples The following sample implements `IChatClient` to show the general structure. For more realistic, concrete implementations of `IChatClient`, see: - [OpenAIChatClient.cs](https://github.com/dotnet/extensions/blob/main/src/Libraries/Microsoft.Extensions.AI.OpenAI/OpenAIChatClient.cs) - [Microsoft.Extensions.AI chat clients](https://github.com/dotnet/extensions/tree/main/src/Libraries/Microsoft.Extensions.AI/ChatCompletion) ## Chat reduction (experimental) > [!IMPORTANT] > This feature is experimental and subject to change. Chat reduction helps manage conversation history by limiting the number of messages or summarizing older messages when the conversation exceeds a specified length. The `Microsoft.Extensions.AI` library provides reducers like <xref:Microsoft.Extensions.AI.MessageCountingChatReducer> that limits the number of non-system messages, and <xref:Microsoft.Extensions.AI.SummarizingChatReducer> that automatically summarizes older messages while preserving context. -
iembeddinggenerator.md 4.2 KB
--- title: Use the IEmbeddingGenerator interface description: Learn how to use the IEmbeddingGenerator interface to generate embeddings for a collection of input values, with optional configuration and cancellation support. ms.date: 12/11/2025 no-loc: ["IEmbeddingGenerator"] --- # Use the IEmbeddingGenerator interface The <xref:Microsoft.Extensions.AI.IEmbeddingGenerator`2> interface represents a generic generator of embeddings. For the generic type parameters, `TInput` is the type of input values being embedded, and `TEmbedding` is the type of generated embedding, which inherits from the <xref:Microsoft.Extensions.AI.Embedding> class. The `Embedding` class serves as a base class for embeddings generated by an `IEmbeddingGenerator`. It's designed to store and manage the metadata and data associated with embeddings. Derived types, like <xref:Microsoft.Extensions.AI.Embedding`1>, provide the concrete embedding vector data. For example, an `Embedding<float>` exposes a `ReadOnlyMemory<float> Vector { get; }` property for access to its embedding data. The `IEmbeddingGenerator` interface defines a method to asynchronously generate embeddings for a collection of input values, with optional configuration and cancellation support. It also provides metadata describing the generator and allows for the retrieval of strongly typed services that can be provided by the generator or its underlying services. ## Create embeddings The primary operation performed with an <xref:Microsoft.Extensions.AI.IEmbeddingGenerator`2> is embedding generation, which is accomplished with its <xref:Microsoft.Extensions.AI.IEmbeddingGenerator`2.GenerateAsync*> method. Accelerator extension methods also exist to simplify common cases, such as generating an embedding vector from a single input. ## Pipelines of functionality As with `IChatClient`, `IEmbeddingGenerator` implementations can be layered. `Microsoft.Extensions.AI` provides a delegating implementation for `IEmbeddingGenerator` for caching and telemetry. The `IEmbeddingGenerator` enables building custom middleware that extends the functionality of an `IEmbeddingGenerator`. The <xref:Microsoft.Extensions.AI.DelegatingEmbeddingGenerator`2> class is an implementation of the `IEmbeddingGenerator<TInput, TEmbedding>` interface that serves as a base class for creating embedding generators that delegate their operations to another `IEmbeddingGenerator<TInput, TEmbedding>` instance. It allows for chaining multiple generators in any order, passing calls through to an underlying generator. The class provides default implementations for methods such as <xref:Microsoft.Extensions.AI.DelegatingEmbeddingGenerator`2.GenerateAsync*> and `Dispose`, which forward the calls to the inner generator instance, enabling flexible and modular embedding generation. The following is an example implementation of such a delegating embedding generator that rate-limits embedding generation requests: This can then be layered around an arbitrary `IEmbeddingGenerator<string, Embedding<float>>` to rate limit all embedding generation operations. In this way, the `RateLimitingEmbeddingGenerator` can be composed with other `IEmbeddingGenerator<string, Embedding<float>>` instances to provide rate-limiting functionality. ## Implementation examples Most users don't need to implement the `IEmbeddingGenerator` interface. However, if you're a library author, then it might be helpful to look at these implementation examples. The following code shows how the `SampleEmbeddingGenerator` class implements the `IEmbeddingGenerator<TInput,TEmbedding>` interface. It has a primary constructor that accepts an endpoint and model ID, which are used to identify the generator. It also implements the <xref:Microsoft.Extensions.AI.IEmbeddingGenerator`2.GenerateAsync(System.Collections.Generic.IEnumerable{`0},Microsoft.Extensions.AI.EmbeddingGenerationOptions,System.Threading.CancellationToken)> method to generate embeddings for a collection of input values. This sample implementation just generates random embedding vectors. For a more realistic, concrete implementation, see [OpenTelemetryEmbeddingGenerator.cs](https://github.com/dotnet/extensions/blob/main/src/Libraries/Microsoft.Extensions.AI/Embeddings/OpenTelemetryEmbeddingGenerator.cs). -
microsoft-extensions-ai.md 6.4 KB
--- title: Microsoft.Extensions.AI libraries description: Learn how to use the Microsoft.Extensions.AI libraries to integrate and interact with various AI services in your .NET applications. ms.date: 12/10/2025 --- # Microsoft.Extensions.AI libraries .NET developers need to integrate and interact with a growing variety of artificial intelligence (AI) services in their apps. The `Microsoft.Extensions.AI` libraries provide a unified approach for representing generative AI components, and enable seamless integration and interoperability with various AI services. This article introduces the libraries and provides in-depth usage examples to help you get started. ## The packages The [📦 Microsoft.Extensions.AI.Abstractions](https://www.nuget.org/packages/Microsoft.Extensions.AI.Abstractions) package provides the core exchange types, including <xref:Microsoft.Extensions.AI.IChatClient> and <xref:Microsoft.Extensions.AI.IEmbeddingGenerator`2>. Any .NET library that provides an LLM client can implement the `IChatClient` interface to enable seamless integration with consuming code. The [📦 Microsoft.Extensions.AI](https://www.nuget.org/packages/Microsoft.Extensions.AI) package has an implicit dependency on the `Microsoft.Extensions.AI.Abstractions` package. This package enables you to easily integrate components such as automatic function tool invocation, telemetry, and caching into your applications using familiar dependency injection and middleware patterns. For example, it provides the <xref:Microsoft.Extensions.AI.OpenTelemetryChatClientBuilderExtensions.UseOpenTelemetry(Microsoft.Extensions.AI.ChatClientBuilder,Microsoft.Extensions.Logging.ILoggerFactory,System.String,System.Action{Microsoft.Extensions.AI.OpenTelemetryChatClient})> extension method, which adds OpenTelemetry support to the chat client pipeline. ### Which package to reference To have access to higher-level utilities for working with generative AI components, reference the `Microsoft.Extensions.AI` package instead (which itself references `Microsoft.Extensions.AI.Abstractions`). Most consuming applications and services should reference the `Microsoft.Extensions.AI` package along with one or more libraries that provide concrete implementations of the abstractions. Libraries that provide implementations of the abstractions typically reference only `Microsoft.Extensions.AI.Abstractions`. ### Install the packages For information about how to install NuGet packages, see [dotnet package add](../core/tools/dotnet-package-add.md) or [Manage package dependencies in .NET applications](../core/tools/dependencies.md). ## APIs and functionality - [The `IChatClient` interface](#the-ichatclient-interface) - [The `IEmbeddingGenerator` interface](#the-iembeddinggenerator-interface) - [The `IImageGenerator` interface (experimental)](#the-iimagegenerator-interface-experimental) ### The `IChatClient` interface The <xref:Microsoft.Extensions.AI.IChatClient> interface defines a client abstraction responsible for interacting with AI services that provide chat capabilities. It includes methods for sending and receiving messages with multi-modal content (such as text, images, and audio), either as a complete set or streamed incrementally. For more information and detailed usage examples, see [Use the IChatClient interface](ichatclient.md). ### The `IEmbeddingGenerator` interface The <xref:Microsoft.Extensions.AI.IEmbeddingGenerator> interface represents a generic generator of embeddings. For the generic type parameters, `TInput` is the type of input values being embedded, and `TEmbedding` is the type of generated embedding, which inherits from the <xref:Microsoft.Extensions.AI.Embedding> class. For more information and detailed usage examples, see [Use the IEmbeddingGenerator interface](iembeddinggenerator.md). ### The IImageGenerator interface (experimental) The <xref:Microsoft.Extensions.AI.IImageGenerator> interface represents a generator for creating images from text prompts or other input. This interface enables applications to integrate image generation capabilities from various AI services through a consistent API. The interface supports text-to-image generation (by calling <xref:Microsoft.Extensions.AI.IImageGenerator.GenerateAsync(Microsoft.Extensions.AI.ImageGenerationRequest,Microsoft.Extensions.AI.ImageGenerationOptions,System.Threading.CancellationToken)>) and [configuration options](xref:Microsoft.Extensions.AI.ImageGenerationOptions) for image size and format. Like other interfaces in the library, it can be composed with middleware for caching, telemetry, and other cross-cutting concerns. For more information, see [Generate images from text using AI](quickstarts/text-to-image.md). ## Build with Microsoft.Extensions.AI You can start building with `Microsoft.Extensions.AI` in the following ways: - **Library developers**: If you own libraries that provide clients for AI services, consider implementing the interfaces in your libraries. This allows users to easily integrate your NuGet package via the abstractions. For examples, see [IChatClient implementation examples](ichatclient.md#implementation-examples) and [IEmbeddingGenerator implementation examples](iembeddinggenerator.md#implementation-examples). - **Service consumers**: If you're developing libraries that consume AI services, use the abstractions instead of hardcoding to a specific AI service. This approach gives your consumers the flexibility to choose their preferred provider. - **Application developers**: Use the abstractions to simplify integration into your apps. This enables portability across models and services, facilitates testing and mocking, leverages middleware provided by the ecosystem, and maintains a consistent API throughout your app, even if you use different services in different parts of your application. - **Ecosystem contributors**: If you're interested in contributing to the ecosystem, consider writing custom middleware components. For more samples, see the [dotnet/ai-samples](https://aka.ms/meai-samples) GitHub repository. For an end-to-end sample, see [eShopSupport](https://github.com/dotnet/eShopSupport). ## See also - [Request a response with structured output](./quickstarts/structured-output.md) - [Build an AI chat app with .NET](./quickstarts/build-chat-app.md) - [Dependency injection in .NET](../core/extensions/dependency-injection/overview.md) - [Caching in .NET](../core/extensions/caching.md) - [Rate limit an HTTP handler in .NET](../core/extensions/http-ratelimiter.md) -
overview.md 3.9 KB
--- title: Develop .NET apps with AI features description: Learn how you can build .NET applications that include AI features. ms.date: 12/10/2025 ms.topic: overview --- # Develop .NET apps with AI features With .NET, you can use artificial intelligence (AI) to automate and accomplish complex tasks in your applications using the tools, platforms, and services that are familiar to you. ## Why choose .NET to build AI apps? Millions of developers use .NET to create applications that run on the web, on mobile and desktop devices, or in the cloud. By using .NET to integrate AI into your applications, you can take advantage of all that .NET has to offer: * A unified story for building web UIs, APIs, and applications. * Supported on Windows, macOS, and Linux. * Is open-source and community-focused. * Runs on top of the most popular web servers and cloud platforms. * Provides powerful tooling to edit, debug, test, and deploy. ## Supported AI providers .NET libraries support a wide range of AI service providers, enabling you to build applications with the AI platform that best fits your needs. The following table lists the major AI providers that integrate with `Microsoft.Extensions.AI`: | Provider | Description | |----------|------------------------|---------------------------|-----------------|-------------| | OpenAI | Direct integration with OpenAI's models including GPT-4, GPT-3.5, and DALL-E | | Azure OpenAI | Enterprise-grade OpenAI models hosted on Azure with enhanced security and compliance | | Azure AI Foundry | Microsoft's managed platform for building and deploying AI agents at scale | | GitHub Models | Access to models available through GitHub's AI model marketplace | | Ollama | Run open-source models locally, for example, Llama, Mistral, and Phi-3 | | Google Gemini | Google's multimodal AI models | | Amazon Bedrock | AWS's managed service for foundation models | Any AI provider that's usable with `Microsoft.Extensions.AI` is also usable with Agent Framework. ## What can you build with AI and .NET? The opportunities with AI are near endless. Here are a few examples of solutions you can build using AI in your .NET applications: * Language processing: Create virtual agents or chatbots to talk with your data and generate content and images. * Computer vision: Identify objects in an image or video. * Audio generation: Use synthesized voices to interact with customers. * Classification: Label the severity of a customer-reported issue. * Task automation: Automatically perform the next step in a workflow as tasks are completed. ## Recommended learning path We recommend the following sequence of tutorials and articles for an introduction to developing applications with AI and .NET: | Scenario | Tutorial | |-----------------------------|-------------------------------------------------------------------------| | Create a chat application | [Build an Azure AI chat app with .NET](./quickstarts/build-chat-app.md) | | Summarize text | [Summarize text using Azure AI chat app](./quickstarts/prompt-model.md) | | Chat with your data | [Get insight about your data from a .NET Azure AI chat app](./quickstarts/build-vector-search-app.md) | | Call .NET functions with AI | [Extend Azure AI using tools and execute a local function with .NET](./quickstarts/use-function-calling.md) | | Generate images | [Generate images from text](./quickstarts/text-to-image.md) | | Train your own model | [ML.NET tutorial](https://dotnet.microsoft.com/learn/ml-dotnet/get-started-tutorial/intro) | Browse the table of contents to learn more about the core concepts, starting with [How generative AI and LLMs work](./conceptual/how-genai-and-llms-work.md). ## Next steps * [Quickstart: Build an Azure AI chat app with .NET](./quickstarts/build-chat-app.md) * [Video series: Machine Learning and AI with .NET](/shows/machine-learning-and-ai-with-dotnet-for-beginners)
-
-
evaluation.md 3.7 KB
# Microsoft.Extensions.AI Evaluation ## Package Set | Package | Purpose | |---|---| | `Microsoft.Extensions.AI.Evaluation` | Core evaluation abstractions and result types | | `Microsoft.Extensions.AI.Evaluation.Quality` | LLM-based quality evaluators such as relevance, completeness, groundedness, and fluency | | `Microsoft.Extensions.AI.Evaluation.NLP` | Non-LLM text-similarity evaluators such as BLEU, GLEU, and F1 | | `Microsoft.Extensions.AI.Evaluation.Safety` | Safety evaluators backed by the Microsoft Foundry Evaluation service | | `Microsoft.Extensions.AI.Evaluation.Reporting` | Result storage, cached responses, and report generation | | `Microsoft.Extensions.AI.Evaluation.Reporting.Azure` | Azure Storage-backed reporting and caching support | | `Microsoft.Extensions.AI.Evaluation.Console` | `dotnet aieval` CLI for reports and cache management | ## Choose Evaluators By Risk ### Quality Use these when answer quality or agent behavior matters: - `RelevanceEvaluator` - `CompletenessEvaluator` - `RetrievalEvaluator` - `FluencyEvaluator` - `CoherenceEvaluator` - `EquivalenceEvaluator` - `GroundednessEvaluator` - `IntentResolutionEvaluator` - `TaskAdherenceEvaluator` - `ToolCallAccuracyEvaluator` ### NLP Use these when you already have reference answers and need cheaper deterministic comparisons: - `BLEUEvaluator` - `GLEUEvaluator` - `F1Evaluator` ### Safety Use these when harmful output, prompt attacks, or unsafe code are part of the release risk: - `ContentHarmEvaluator` - `ProtectedMaterialEvaluator` - `GroundednessProEvaluator` - `UngroundedAttributesEvaluator` - `HateAndUnfairnessEvaluator` - `SelfHarmEvaluator` - `ViolenceEvaluator` - `SexualEvaluator` - `CodeVulnerabilityEvaluator` - `IndirectAttackEvaluator` ## Practical Evaluation Loop 1. Pick a stable prompt or scenario set that represents the real feature. 2. Decide whether the gate is about answer quality, tool behavior, safety, or all three. 3. Use the same `IChatClient`-backed app surface that production uses, or a controlled test double when you are isolating logic. 4. Cache responses for repeatability and lower cost. 5. Store results and publish reports so model, prompt, or middleware changes are comparable across runs. ## CI Guidance - Use NLP evaluators for low-cost baseline checks on every PR when reference outputs exist. - Use quality evaluators on targeted, high-value scenarios such as retrieval, summarization, tool use, or task adherence. - Use safety evaluators for user-facing or code-producing features before release. - Track threshold changes deliberately; do not quietly relax gates when a prompt or model regresses. ## Agent-Oriented Checks Even if the app is not using full Agent Framework, agent-like workflows often need: - `IntentResolutionEvaluator` when the system has to understand and complete multi-step user requests - `TaskAdherenceEvaluator` when the system receives bounded instructions or policies - `ToolCallAccuracyEvaluator` when local functions or MCP-backed tools are part of the flow These metrics are often the first place where prompt drift or tool-schema changes show up. ## Reporting And Caching - The libraries support response caching so unchanged prompt-model combinations can reuse prior results. - Reporting packages let you persist evaluation data and generate human-readable reports. - The `dotnet aieval` CLI is useful for report generation and cache management in local runs or CI pipelines. ## Common Failure Modes - Evaluating only one happy-path prompt instead of the real scenario envelope. - Comparing outputs without fixing the prompt, grounding data, or model selection. - Treating evaluation as a one-time benchmark instead of a regression suite. - Shipping tool-using or RAG features without measuring task adherence, groundedness, or tool accuracy. -
examples.md 4.6 KB
# Microsoft.Extensions.AI Practical Examples ## Quickstart-To-Task Map | Scenario | Start with | Main packages or surfaces | Notes | |---|---|---|---| | Prompt a model once | `official-docs/quickstarts/prompt-model.md` | `Microsoft.Extensions.AI.OpenAI` + provider SDK | Smallest provider-agnostic entry point | | Build a chat app | `official-docs/quickstarts/build-chat-app.md` | `IChatClient` | Good baseline for message history and follow-up turns | | Stream responses in UI | `official-docs/ichatclient.md` | `GetStreamingResponseAsync` | Use `IAsyncEnumerable<ChatResponseUpdate>` all the way to the UI | | Request structured output | `official-docs/quickstarts/structured-output.md` | typed `GetResponseAsync<T>` helpers | Prefer typed enums or records over manual JSON parsing | | Execute local tools | `official-docs/quickstarts/use-function-calling.md` | `AIFunction`, `FunctionInvokingChatClient` | Add invalid-input handling from `official-docs/how-to/handle-invalid-tool-input.md` | | Build vector search or RAG | `official-docs/quickstarts/build-vector-search-app.md` | `IEmbeddingGenerator`, `Microsoft.Extensions.VectorData.Abstractions` | Keep chunking and embedding model/version stable | | Process data for RAG | `official-docs/quickstarts/process-data.md` | `Microsoft.Extensions.DataIngestion`, `IngestionPipeline<T>` | Use when the ingestion pipeline matters as much as inference | | Chat with a local model | `official-docs/quickstarts/chat-local-model.md` | local provider adapter + `IChatClient` | Good for dev, lower cost, and offline workflows | | Generate images | `official-docs/quickstarts/text-to-image.md` | experimental `IImageGenerator` or provider client | Treat image generation as a separate capability surface | | Build an MCP client | `official-docs/quickstarts/build-mcp-client.md` | MCP client + `IChatClient` | Relevant when tools live behind MCP servers | | Build an MCP server | `official-docs/quickstarts/build-mcp-server.md` | MCP server SDK | This leans toward `dotnet-mcp`, but often pairs with Extensions.AI clients | | Create a minimal assistant | `official-docs/quickstarts/create-assistant.md` | provider-specific assistants SDK | This quickstart is assistant-service-centric, not the pure `IChatClient` abstraction layer | ## Recommended Composition Recipes ### Provider-Agnostic App - Register one or more `IChatClient` implementations in DI. - Add options configuration, logging or telemetry, caching, and function invocation in a deliberate builder order. - Keep feature code dependent on `IChatClient`, not the vendor SDK, unless you truly need provider-specific capabilities. ### Typed Chat + Tools - Use `GetResponseAsync<T>` or the equivalent typed helpers for structured output. - Give the model a narrow result shape and a narrow tool surface. - Route ambient tool data through `AdditionalProperties`, `AIFunctionArguments`, or DI instead of serializing hidden state into prompts. ### Vector Search / RAG - Use `IEmbeddingGenerator<string, Embedding<float>>` to create embeddings for both source content and user queries. - Store vectors in a vector store accessed through `Microsoft.Extensions.VectorData.Abstractions`. - Keep ingestion, chunking, and retrieval policies versioned so evaluation results stay meaningful over time. ### Data Ingestion for RAG - Start from `Microsoft.Extensions.DataIngestion` when documents must be read, normalized, enriched, chunked, and written as one pipeline instead of a pile of custom helpers. - Reach for the official processing shape: - document reader such as MarkItDown or Markdig - optional document processor such as `ImageAlternativeTextEnricher` - chunker such as `HeaderChunker` or semantic chunking - chunk processors such as `SummaryEnricher` - `VectorStoreWriter<T>` and `IngestionPipeline<T>` for the final persisted flow - Handle `ProcessAsync` results per document. A single ingestion failure should be an explicit policy decision, not an accidental crash. ### Evaluation-Backed Delivery - Add quality and safety evaluators for important prompts and user journeys. - Run cheap NLP evaluators for stable offline comparisons when you have reference outputs. - Publish reports and reuse cached evaluation responses in CI so the team can compare prompt or model changes. ## Important Boundaries - `Microsoft.Extensions.AI` is ideal for provider abstraction, middleware, embeddings, evaluation, and typed tool calling. - Provider-hosted assistants APIs are adjacent but not identical to `IChatClient` composition. - When the app needs threads, multi-agent orchestration, or durable workflow control, hand off to `dotnet-microsoft-agent-framework`. -
official-docs-index.md 8.3 KB
# Official Docs Index This skill keeps a slim, markdown-only snapshot of the official `.NET AI` docs tree from `dotnet/docs` under `docs/ai`. ## Snapshot Summary - Local root: `references/official-docs/` - Coverage: `48` useful markdown pages - Scope: `Microsoft.Extensions.AI`, adjacent `VectorData` and `DataIngestion` guidance, evaluation libraries, MCP quickstarts, RAG guidance, and the surrounding `.NET AI` concept pages - Boundary: Microsoft Agent Framework is linked from this docs tree, but its dedicated authored snapshot and deeper routing guidance live in the separate `dotnet-microsoft-agent-framework` skill - Intentional exclusions: snippet trees, project files, TOC scaffolding, DocFX support files, JSON helpers, media folders, and other low-signal assets are not mirrored into the skill ## Start Here - [`official-docs/overview.md`](official-docs/overview.md) - Root `.NET AI` landing page - [`official-docs/dotnet-ai-ecosystem.md`](official-docs/dotnet-ai-ecosystem.md) - Ecosystem map and the official boundary between `Microsoft.Extensions.AI` and Agent Framework - [`official-docs/microsoft-extensions-ai.md`](official-docs/microsoft-extensions-ai.md) - Package split and core API overview - [`official-docs/ichatclient.md`](official-docs/ichatclient.md) - Chat, streaming, tools, caching, telemetry, DI, and state handling - [`official-docs/iembeddinggenerator.md`](official-docs/iembeddinggenerator.md) - Embeddings, delegating generators, and implementation guidance ## Section Map - Root pages: `overview.md`, `dotnet-ai-ecosystem.md`, `microsoft-extensions-ai.md`, `ichatclient.md`, `iembeddinggenerator.md`, `get-started-mcp.md`, `get-started-app-chat-template.md`, `get-started-app-chat-scaling-with-azure-container-apps.md`, `azure-ai-services-authentication.md` - Concepts: [`official-docs/conceptual/`](official-docs/conceptual/) with `11` pages covering agents, tools, tokens, embeddings, vector databases, ingestion, prompt engineering, zero-shot and few-shot, chain-of-thought, and RAG - Quickstarts: [`official-docs/quickstarts/`](official-docs/quickstarts/) with `14` pages covering prompting, chat apps, structured output, vector search, function calling, local models, assistants, MCP client and server, templates, text-to-image, and data processing - How-to: [`official-docs/how-to/`](official-docs/how-to/) with `5` pages covering function data access, invalid tool input, content filtering, Azure-hosted auth, and tokenizers - Evaluation: [`official-docs/evaluation/`](official-docs/evaluation/) with `5` pages covering responsible AI, libraries, response quality, reporting, and safety evaluation - Resources: [`official-docs/resources/`](official-docs/resources/) with `3` pages for general `.NET AI`, Azure AI, and MCP resource lists - Tutorial: [`official-docs/tutorials/tutorial-ai-vector-search.md`](official-docs/tutorials/tutorial-ai-vector-search.md) for the deeper vector-search walkthrough ## Complete Local File Map ### Root Pages - [`official-docs/azure-ai-services-authentication.md`](official-docs/azure-ai-services-authentication.md) - [`official-docs/dotnet-ai-ecosystem.md`](official-docs/dotnet-ai-ecosystem.md) - [`official-docs/get-started-app-chat-scaling-with-azure-container-apps.md`](official-docs/get-started-app-chat-scaling-with-azure-container-apps.md) - [`official-docs/get-started-app-chat-template.md`](official-docs/get-started-app-chat-template.md) - [`official-docs/get-started-mcp.md`](official-docs/get-started-mcp.md) - [`official-docs/ichatclient.md`](official-docs/ichatclient.md) - [`official-docs/iembeddinggenerator.md`](official-docs/iembeddinggenerator.md) - [`official-docs/microsoft-extensions-ai.md`](official-docs/microsoft-extensions-ai.md) - [`official-docs/overview.md`](official-docs/overview.md) ### Conceptual - [`official-docs/conceptual/agents.md`](official-docs/conceptual/agents.md) - [`official-docs/conceptual/ai-tools.md`](official-docs/conceptual/ai-tools.md) - [`official-docs/conceptual/chain-of-thought-prompting.md`](official-docs/conceptual/chain-of-thought-prompting.md) - [`official-docs/conceptual/data-ingestion.md`](official-docs/conceptual/data-ingestion.md) - [`official-docs/conceptual/embeddings.md`](official-docs/conceptual/embeddings.md) - [`official-docs/conceptual/how-genai-and-llms-work.md`](official-docs/conceptual/how-genai-and-llms-work.md) - [`official-docs/conceptual/prompt-engineering-dotnet.md`](official-docs/conceptual/prompt-engineering-dotnet.md) - [`official-docs/conceptual/rag.md`](official-docs/conceptual/rag.md) - [`official-docs/conceptual/understanding-tokens.md`](official-docs/conceptual/understanding-tokens.md) - [`official-docs/conceptual/vector-databases.md`](official-docs/conceptual/vector-databases.md) - [`official-docs/conceptual/zero-shot-learning.md`](official-docs/conceptual/zero-shot-learning.md) ### How-To - [`official-docs/how-to/access-data-in-functions.md`](official-docs/how-to/access-data-in-functions.md) - [`official-docs/how-to/app-service-aoai-auth.md`](official-docs/how-to/app-service-aoai-auth.md) - [`official-docs/how-to/content-filtering.md`](official-docs/how-to/content-filtering.md) - [`official-docs/how-to/handle-invalid-tool-input.md`](official-docs/how-to/handle-invalid-tool-input.md) - [`official-docs/how-to/use-tokenizers.md`](official-docs/how-to/use-tokenizers.md) ### Quickstarts - [`official-docs/quickstarts/ai-templates.md`](official-docs/quickstarts/ai-templates.md) - [`official-docs/quickstarts/build-chat-app.md`](official-docs/quickstarts/build-chat-app.md) - [`official-docs/quickstarts/build-mcp-client.md`](official-docs/quickstarts/build-mcp-client.md) - [`official-docs/quickstarts/build-mcp-server.md`](official-docs/quickstarts/build-mcp-server.md) - [`official-docs/quickstarts/build-vector-search-app.md`](official-docs/quickstarts/build-vector-search-app.md) - [`official-docs/quickstarts/chat-local-model.md`](official-docs/quickstarts/chat-local-model.md) - [`official-docs/quickstarts/create-assistant.md`](official-docs/quickstarts/create-assistant.md) - [`official-docs/quickstarts/generate-images.md`](official-docs/quickstarts/generate-images.md) - [`official-docs/quickstarts/process-data.md`](official-docs/quickstarts/process-data.md) - [`official-docs/quickstarts/prompt-model.md`](official-docs/quickstarts/prompt-model.md) - [`official-docs/quickstarts/publish-mcp-registry.md`](official-docs/quickstarts/publish-mcp-registry.md) - [`official-docs/quickstarts/structured-output.md`](official-docs/quickstarts/structured-output.md) - [`official-docs/quickstarts/text-to-image.md`](official-docs/quickstarts/text-to-image.md) - [`official-docs/quickstarts/use-function-calling.md`](official-docs/quickstarts/use-function-calling.md) ### Evaluation - [`official-docs/evaluation/evaluate-ai-response.md`](official-docs/evaluation/evaluate-ai-response.md) - [`official-docs/evaluation/evaluate-safety.md`](official-docs/evaluation/evaluate-safety.md) - [`official-docs/evaluation/evaluate-with-reporting.md`](official-docs/evaluation/evaluate-with-reporting.md) - [`official-docs/evaluation/libraries.md`](official-docs/evaluation/libraries.md) - [`official-docs/evaluation/responsible-ai.md`](official-docs/evaluation/responsible-ai.md) ### Resources - [`official-docs/resources/azure-ai.md`](official-docs/resources/azure-ai.md) - [`official-docs/resources/get-started.md`](official-docs/resources/get-started.md) - [`official-docs/resources/mcp-servers.md`](official-docs/resources/mcp-servers.md) ### Tutorials - [`official-docs/tutorials/tutorial-ai-vector-search.md`](official-docs/tutorials/tutorial-ai-vector-search.md) ## API Reference Landing Pages - `https://learn.microsoft.com/dotnet/api/microsoft.extensions.ai` - `https://learn.microsoft.com/dotnet/api/microsoft.extensions.vectordata` - `https://learn.microsoft.com/dotnet/api/microsoft.extensions.dataingestion` ## Reading Strategy - Use the local snapshot when exact wording, package names, or Learn-page structure matters. - Start with the authored overview pages before diving into provider-specific quickstarts. - Raw Learn `:::code` and `:::image` source-asset directives are stripped from the local snapshot to keep it prose-first and avoid broken local references. - For orchestration, threads, workflows, or hosted-agent protocols, switch to the `dotnet-microsoft-agent-framework` skill rather than assuming the answer lives in the `Microsoft.Extensions.AI` layer. -
patterns.md 5.2 KB
# Microsoft.Extensions.AI Patterns ## Package Selection - Applications and services should usually reference `Microsoft.Extensions.AI`. - Provider libraries and reusable connectors should usually reference `Microsoft.Extensions.AI.Abstractions`. - Add `Microsoft.Extensions.VectorData.Abstractions` when you need vector-store CRUD or search. - Add `Microsoft.Extensions.DataIngestion` when you need document ingestion and preparation for RAG. - Add `Microsoft.Extensions.AI.Evaluation.*` when prompts, tool use, or safety need measurable regression checks. ## `IChatClient` Request Model - Use `GetResponseAsync` for whole responses and `GetStreamingResponseAsync` for UI or interactive console streaming. - Treat `ChatResponse.Messages` as the provider-independent response payload. Do not assume a single plain-text string is the only output. - Use `ChatOptions` for model selection, temperature, tools, additional provider properties, and raw provider-specific options. - For stateless providers, replay the relevant message history each turn. - For stateful providers, propagate `ConversationId` from `ChatResponse` to `ChatOptions` instead of manually resending all prior turns. ## Middleware and Builder Composition - Build chat pipelines explicitly with `ChatClientBuilder`. - Keep middleware order deliberate. A common pattern is: configure options, add logging or telemetry, add caching, then add function invocation. - Use keyed DI registrations when the app needs multiple chat clients or multiple model classes. - Prefer cross-cutting middleware over ad-hoc wrappers scattered through feature code. ## Tool Calling - Describe tools with `AIFunction` and `AIFunctionFactory`. - Use `FunctionInvokingChatClient` when you want automatic tool invocation instead of manually inspecting tool-related message content. - Keep tool registration scoped to the current conversation or job. Tool descriptions count against token limits and can become a hidden cost/latency tax when the list grows. - Remember that tools can be local .NET methods, external APIs, or MCP-backed operations. The model chooses; your application still validates arguments and decides whether the call is safe to execute. - Pass ambient data through closures, `ChatOptions.AdditionalProperties`, `AIFunctionArguments.Context`, or DI, depending on lifetime and ownership. - Validate invalid tool input explicitly. Do not trust the model to always produce perfectly shaped arguments. - Keep side effects narrow, auditable, and guarded outside the prompt. ## Data Ingestion Pipelines - Use `Microsoft.Extensions.DataIngestion` when the document-processing side of RAG matters as much as retrieval or prompting. - Model the flow explicitly: document reader -> document processor -> chunker -> chunk processor -> writer. - Start from `IngestionDocument` and choose built-in readers like MarkItDown or Markdig instead of ad-hoc file-to-string helpers when document fidelity matters. - Use `ImageAlternativeTextEnricher` at the document level and chunk enrichers like `SummaryEnricher`, `KeywordEnricher`, or `ClassificationEnricher` at the chunk level. - Keep tokenizer, chunk size, overlap, and embedding model choices versioned together so reindexing and evaluation stay comparable. - Treat `IngestionPipeline<T>.ProcessAsync` as partial success. Handle `IngestionResult` per document and decide explicitly whether to retry, continue, or fail the batch. ## Structured Output - Use typed response helpers when you need enums, records, or other constrained result shapes. - Keep requested schemas small and stable. The more ambiguous the target type, the more fragile the model output becomes. - Log raw provider output or failure details when typed deserialization fails. ## Embeddings and Vector Search - Use `IEmbeddingGenerator<TInput, TEmbedding>` for semantic indexing, similarity, search, and embedding-backed caches. - Keep the embedding model fixed per collection or explicitly versioned. Mixing models in one vector space causes silent quality degradation. - Ensure vector-store dimensions match the embedding model output. - Keep chunking deterministic so reindexing and evaluation remain reproducible. - Use delegating generators or wrappers for telemetry, rate limits, and caching rather than duplicating those concerns at each call site. ## Evaluation - Use quality evaluators when answer relevance, completeness, truthfulness, or groundedness matter. - Use agent-focused evaluators like `IntentResolutionEvaluator`, `TaskAdherenceEvaluator`, and `ToolCallAccuracyEvaluator` when workflows depend on tool use or instruction following. - Use NLP evaluators for cheaper offline regression baselines when you already have reference answers. - Use reporting and response caching in CI so evaluation runs are reproducible and affordable. ## Escalate to Agent Framework When - the application needs agent threads or durable interaction state - the control flow becomes multi-step, multi-agent, or workflow-driven - remote hosting protocols, A2A, AG-UI, or durable execution enter the design - the architecture needs more than model abstraction and middleware composition `Microsoft.Extensions.AI` is the composition layer. `Microsoft Agent Framework` is the orchestration layer built on top of the abstractions.
-
-
SKILL.md 11 KB
--- name: dotnet-microsoft-extensions-ai version: "1.3.0" category: "AI" description: "Build provider-agnostic .NET AI integrations with `Microsoft.Extensions.AI`, `IChatClient`, embeddings, middleware, structured output, vector search, and evaluation." compatibility: "Requires `Microsoft.Extensions.AI` or a .NET AI application that needs model, embedding, tool-calling, or evaluation composition without full agent orchestration." --- # Microsoft.Extensions.AI ## Trigger On - building or reviewing `.NET` code that uses `Microsoft.Extensions.AI`, `Microsoft.Extensions.AI.Abstractions`, `IChatClient`, `IEmbeddingGenerator`, `ChatOptions`, or `AIFunction` - adding `IImageGenerator`, local-model chat via Ollama, AI app templates, or the `.NET AI` quickstarts for assistants and MCP - choosing between low-level AI abstractions, provider SDKs, vector-search composition, evaluation libraries, and a fuller agent framework - adding streaming chat, structured output, embeddings, tool calling, telemetry, caching, or DI-based AI middleware - wiring `Microsoft.Extensions.VectorData`, `Microsoft.Extensions.DataIngestion`, MCP tooling, or evaluation packages around a provider-agnostic AI app ## Workflow 1. Classify the request first: plain model access, tool calling, embeddings/vector search, evaluation, image generation, local-model prototyping, MCP bootstrap, or true agent orchestration. 2. Default to `Microsoft.Extensions.AI` for application and service code that needs provider-agnostic chat, embeddings, middleware, structured output, and testability. 3. Reference `Microsoft.Extensions.AI.Abstractions` directly only when authoring provider libraries or lower-level reusable integration packages. 4. Model `IChatClient` and `IEmbeddingGenerator` composition explicitly in DI. Keep options, caching, telemetry, logging, and tool invocation inspectable in the pipeline. 5. Treat chat state deliberately. For stateless providers, resend history. For stateful providers, propagate `ConversationId` rather than assuming all providers behave the same way. 6. Use `Microsoft.Extensions.VectorData` and `Microsoft.Extensions.DataIngestion` as adjacent building blocks for RAG instead of hand-rolling store abstractions prematurely. Model ingestion as an explicit reader -> processor -> chunker -> writer pipeline when the document-preparation path matters. 7. Treat the `.NET AI` quickstarts as bootstrap paths, not finished architecture. They now cover minimal assistants, MCP client/server flows, local models, app templates, and image generation. Start there for a vertical slice, then harden the DI, telemetry, and evaluation story here. 8. Escalate to `dotnet-microsoft-agent-framework` when the requirement becomes agent threads, multi-agent orchestration, higher-order workflows, durable execution, or remote agent hosting. 9. Validate with real providers, realistic prompts, and evaluation gates so the abstraction layer actually buys portability and reliability. ## Architecture ```mermaid flowchart LR A["Task"] --> B{"Need agent threads, multi-agent orchestration, or remote agent hosting?"} B -->|Yes| C["Use Microsoft Agent Framework on top of `Microsoft.Extensions.AI.Abstractions`"] B -->|No| D{"Need provider-agnostic chat, embeddings, tools, typed output, or evaluation?"} D -->|Yes| E["Use `Microsoft.Extensions.AI`"] E --> F["Compose `IChatClient` / `IEmbeddingGenerator` in DI"] F --> G["Add caching, telemetry, tools, vector data, and evaluation deliberately"] D -->|No| H["Use plain provider SDKs or deterministic .NET code"] ``` ## Core Knowledge - `Microsoft.Extensions.AI.Abstractions` contains the core exchange contracts such as `IChatClient`, `IEmbeddingGenerator<TInput, TEmbedding>`, message/content types, and tool abstractions. - `Microsoft.Extensions.AI` adds the higher-level application surface: middleware builders, automatic function invocation, caching, logging, and OpenTelemetry integration. - Most apps and services should reference `Microsoft.Extensions.AI`; provider and connector libraries usually reference only the abstractions package. - `IChatClient` centers on `GetResponseAsync` and `GetStreamingResponseAsync`. The returned `ChatResponse` or `ChatResponseUpdate` objects carry messages, tool-related content, metadata, and optional conversation identifiers. - Local-model quickstarts still route through the same `IChatClient` abstraction. Ollama-backed clients are useful for low-cost prototyping, offline dev loops, and portability testing, but you still own chat history replay, latency, and model-quality tradeoffs. - `ChatOptions` is the normal control plane for model ID, temperature, tools, `AdditionalProperties`, and provider-specific raw options. - Tool calling is modeled with `AIFunction`, `AIFunctionFactory`, and `FunctionInvokingChatClient`. Ambient data can flow through closures, `AdditionalProperties`, `AIFunctionArguments.Context`, or DI. - Tool calling can target local .NET methods, external APIs, or MCP-backed tools. The model requests calls; your app still owns execution, validation, and side-effect boundaries. - Tool definitions consume request tokens. Keep tool descriptions short and register only the tools relevant for the current conversation or workflow. - `FunctionInvokingChatClient` can handle the tool-invocation loop and parallel tool-call responses automatically when the provider/model supports that shape. - `IEmbeddingGenerator` is the standard abstraction for semantic search, vector indexing, similarity, and cache-key generation. Pair it with `Microsoft.Extensions.VectorData.Abstractions` for vector store operations. - `IImageGenerator` is the experimental MEAI image surface. Treat `MEAI001` as an intentional opt-in, keep image generation separate from chat concerns, and compose logging/caching/hosting middleware around it the same way you would for `IChatClient`. - `Microsoft.Extensions.DataIngestion` gives you the document-side RAG pipeline: `IngestionDocument`, document readers like MarkItDown/Markdig, document processors such as `ImageAlternativeTextEnricher`, chunkers, chunk processors, `VectorStoreWriter<T>`, and `IngestionPipeline<T>` for end-to-end composition. - `IngestionPipeline<T>.ProcessAsync` is partial-success oriented. Handle `IAsyncEnumerable<IngestionResult>` deliberately instead of assuming one failed document should automatically crash the whole ingestion run. - `Microsoft.Extensions.AI.Evaluation.*` gives you quality, NLP, safety, caching, and reporting layers for regression checks and CI gates. - The official `.NET AI` docs now make MCP, assistants, local models, templates, and text-to-image part of the same app-level story. Use `dotnet-mcp` when the protocol itself becomes the design problem; stay here when you still mostly need app composition around `IChatClient` and friends. - `Microsoft Agent Framework` builds on these abstractions. Use it when you need autonomous orchestration, threads, workflows, hosting, or multi-agent collaboration instead of just model composition. ## Decision Cheatsheet | If you need | Default choice | Why | |---|---|---| | App-level provider abstraction with middleware | `Microsoft.Extensions.AI` | Highest leverage for apps and services | | A reusable provider or connector library | `Microsoft.Extensions.AI.Abstractions` | Keeps your package at the contract layer | | Typed chat or UI streaming | `IChatClient` with `GetResponseAsync` / `GetStreamingResponseAsync` | Common request/response shape across providers | | Tool calling from .NET methods | `AIFunction` + `FunctionInvokingChatClient` | Native function metadata and invocation pipeline | | Typed structured output | `IChatClient.GetResponseAsync<T>` extensions | Keeps schema intent in code instead of prompt parsing | | Vector search or RAG | `IEmbeddingGenerator` + `Microsoft.Extensions.VectorData.Abstractions` | Standardizes embeddings and store access | | Local model prototyping | `IChatClient` with an Ollama-backed implementation | Keeps the app on the MEAI abstractions while you validate prompts or UX locally | | Text-to-image or image-generation middleware | `IImageGenerator` | Use the dedicated image abstraction instead of overloading chat APIs | | Evaluation and regression gates | `Microsoft.Extensions.AI.Evaluation.*` | Relevance, safety, task adherence, caching, reports | | Agent threads or multi-step autonomous orchestration | `dotnet-microsoft-agent-framework` | This is beyond plain provider abstraction | ## Common Failure Modes - Referencing only `Microsoft.Extensions.AI.Abstractions` in an app and then rebuilding middleware, telemetry, or function invocation by hand. - Treating `IChatClient` as if it already gives you durable agent threads, orchestration, or hosted-agent semantics. - Mixing provider-specific assistants APIs with `IChatClient` as if they were the same runtime contract. - Forgetting to distinguish stateless history replay from stateful `ConversationId` flows. - Hiding important chat behavior in singleton service fields instead of explicit message history, options, or persistent storage. - Adding tool calling without validating parameter binding, invalid input behavior, side effects, or DI-scoped dependencies. - Building RAG without stable chunking, embedding-model/version tracking, or vector dimension discipline. - Shipping AI features without evaluation baselines, safety checks, or telemetry for prompt/model drift. ## Deliver - a justified package and abstraction choice: `Abstractions` only vs full `Microsoft.Extensions.AI` - a concrete `IChatClient` / `IEmbeddingGenerator` composition strategy - explicit tool-calling, options, state, caching, logging, and telemetry decisions - vector-search, evaluation, or MCP integration guidance when the scenario needs it - a clear escalation path to Agent Framework when the problem exceeds provider abstraction ## Validate - the abstraction layer solves a real portability, testability, or composition problem - provider registration and middleware order stay explicit in DI - chat state management matches whether the provider is stateless or stateful - structured output, tool invocation, and embedding flows are typed and observable - vector store, embedding model, and chunking strategy are consistent - evaluation or safety gates exist for important prompts and agent-like behaviors - agentic requirements are not being under-modeled as a simple `IChatClient` integration When exact wording, edge-case API behavior, or less-common examples matter, check the local official docs snapshot before relying on summaries. ## References - [official-docs-index.md](references/official-docs-index.md) - Slim local snapshot map with direct links to every mirrored `.NET AI` docs page plus API-reference pointers - [patterns.md](references/patterns.md) - Package choice, `IChatClient`, embeddings, DI pipelines, tool-calling, and Agent Framework escalation guidance - [examples.md](references/examples.md) - Quickstart-to-task map covering chat, structured output, function calling, vector search, local models, MCP, and assistants - [evaluation.md](references/evaluation.md) - Quality, NLP, safety, caching, reporting, and CI-oriented evaluation guidance
Comments (0)
Sign in to join the conversation.
Reviews (0)
No reviews yet.
No comments yet.