Large Language Models with .NET

Large Language Models, commonly called LLMs, are the technology behind many modern generative AI applications.

LLMs can process and generate natural language and can also work with code and, depending on the model, multimodal inputs such as images and audio.

For a .NET developer, an LLM can be treated as a service that your application communicates with through an SDK or a common abstraction.

The basic architecture is:

.NET Application
       |
       v
AI Client
       |
       v
LLM API
       |
       v
Large Language Model
       |
       v
Generated Response

Modern .NET provides a common abstraction through Microsoft.Extensions.AI, including IChatClient, while provider SDKs such as the official OpenAI .NET library provide direct access to specific model providers. Microsoft currently positions Microsoft.Extensions.AI as the common application-level abstraction for LLM interaction.

What Is a Large Language Model?

A Large Language Model is a machine-learning model trained on very large amounts of language data to learn statistical patterns in text.

At a high level, an LLM learns relationships between tokens and uses those learned relationships to generate new sequences.

For example:

Input:

Explain dependency injection in C#.

         |
         v

LLM

         |
         v

Generated text

The model does not simply look up an answer from a database.

During generation, it predicts the next token based on the current sequence and the patterns learned during training. Microsoft describes this process as autoregressive generation.

Why Are LLMs Called "Large"?

The word "large" generally refers to the scale of the model and its training process.

An LLM can involve:

Large training datasets
Large neural networks
Large numbers of parameters
Large compute requirements
Large vocabularies
Large context windows

The exact architecture and parameter counts differ between models.

For a .NET developer, the important point is that the model is usually hosted separately from your application.

ASP.NET Core
      |
      v
AI SDK
      |
      v
Model Service
      |
      v
LLM

Your .NET application does not usually need to contain the entire model.

LLMs and Generative AI

Generative AI is the broader category.

It can generate:

Text
Code
Images
Audio
Video
Structured data

LLMs focus primarily on language-related generation and understanding.

Therefore:

Generative AI
      |
      +---- LLMs
      |
      +---- Image Models
      |
      +---- Audio Models
      |
      +---- Multimodal Models

Microsoft's current .NET AI documentation distinguishes generative AI as a broader area and LLMs as the models commonly used for natural-language generation and processing.

How an LLM Works

A simplified LLM request can be visualized as:

User Input
    |
    v
Tokenization
    |
    v
Token IDs
    |
    v
Embeddings / Model Processing
    |
    v
Transformer Layers
    |
    v
Next-Token Probabilities
    |
    v
Selected Token
    |
    v
Repeat
    |
    v
Generated Response

This is a simplified conceptual model, not the complete implementation of a modern LLM.

What Is a Token?

LLMs do not process text exactly as humans see it.

Text is first broken into tokens.

A token can be:

A complete word
A part of a word
Punctuation
A character sequence

For example, a sentence such as:

I am learning C#.

might be represented internally as a sequence of token IDs.

The exact tokenization depends on the model's tokenizer.

Microsoft's documentation explains that tokenization happens before model processing and that tokens can represent words, partial words, character sequences, and punctuation.

Why Tokens Matter in .NET

Tokens matter because AI services often measure usage and context in tokens.

For example:

Input:
    System prompt
    Conversation
    User question
    Retrieved documents

Output:
    Generated answer

Everything consumes part of the model's available context and potentially contributes to usage.

The architecture is:

System Instructions
        +
Conversation
        +
User Input
        +
RAG Context
        +
Tool Results
        |
        v
       Tokens
        |
        v
       Model
        |
        v
   Output Tokens

Tokenization Example

Suppose you have:

Build an ASP.NET Core AI chatbot.

A tokenizer may divide that into several tokens rather than exactly five word tokens.

The actual token sequence is model-dependent.

This is why token estimates such as:

characters / 4

are only rough approximations.

Production systems should use the tokenizer or usage information supported by the specific model/API when accurate accounting matters.

Tokens and Cost

Many AI systems measure usage through:

Input tokens
Output tokens
Total tokens

Conceptually:

Total Tokens
   =
Input Tokens
   +
Output Tokens

If your application sends large amounts of context:

Conversation
+
Documents
+
Tool output

the input token count can increase significantly.

This affects:

Cost
Latency
Context utilization

What Is a Context Window?

A model's context window is the maximum amount of token context it can process for a request.

Conceptually:

+--------------------------------------------+
|              Context Window                |
|                                            |
| System instructions                        |
| Conversation history                      |
| User request                              |
| Retrieved documents                       |
| Tool results                              |
| Generated output                          |
|                                            |
+--------------------------------------------+

Everything must fit within the limits supported by the model and service.

Microsoft's current LLM fundamentals documentation explains that the prompt, conversation history, injected context, and generated tokens all participate in the model's context window.

Why Context Windows Matter

Imagine a chat application where a user has exchanged hundreds of messages.

Sending the entire conversation on every request can create:

Large input
   |
   v
More tokens
   |
   v
Higher latency
   |
   v
Higher usage

This is why applications need:

Conversation trimming
Summarization
Context reduction
Relevant-history retrieval

Microsoft's Microsoft.Extensions.AI documentation currently includes experimental chat-reduction functionality for limiting or summarizing conversation history.

LLM Training vs LLM Inference

Two concepts are important.

Training

The model learns patterns from large datasets.

Training Data
     |
     v
Machine Learning
     |
     v
Model Parameters

Inference

The trained model receives input and generates output.

Prompt
  |
  v
Trained Model
  |
  v
Generated Output

Most .NET developers integrating hosted LLMs are working with inference, not model training.

LLM Application Architecture

A typical .NET application does not train the model.

Instead:

.NET Application
       |
       v
Prompt / Messages
       |
       v
LLM Service
       |
       v
Pretrained Model
       |
       v
Response

This lets developers focus on:

Application logic
Context
Prompts
Data
Tools
Security
User experience
Evaluation

What Is Inference?

Inference is the process where a trained model receives input and produces a prediction or generated output.

For an LLM:

Input tokens
     |
     v
Model computation
     |
     v
Probability distribution
     |
     v
Next token
     |
     v
Repeat

Microsoft's current LLM fundamentals documentation describes this as autoregressive generation, where output is generated one token at a time.

How an LLM Generates Text

Suppose the input is:

C# is a

The model calculates probabilities for possible next tokens.

Conceptually:

professional     0.35
programming      0.30
language         0.20
...

One token is selected.

Then the process repeats:

C# is a
      |
      v
C# is a programming
      |
      v
C# is a programming language
      |
      v
...

The actual probabilities are much more complex than this simplified example.

Why LLM Output Can Differ

Even with the same prompt, output can vary depending on model behavior and sampling configuration.

Factors can include:

Temperature
Sampling configuration
Model version
System instructions
Context
Tool results
Provider behavior

This is why:

Same prompt

does not necessarily guarantee:

Identical output

Temperature

Temperature is a model-generation setting exposed by many model APIs.

Conceptually:

Lower temperature
      |
      v
More constrained sampling

Higher temperature
      |
      v
More variation

It should not be treated as a universal "creativity slider" with identical behavior across all providers and models.

In Microsoft.Extensions.AI, ChatOptions exposes common options such as Temperature.

Example:

ChatOptions options = new()
{
    Temperature = 0.2f
};

ChatResponse response =
    await chatClient.GetResponseAsync(
        "Explain dependency injection in C#.",
        options);

Provider-specific behavior can differ, so test the setting with the actual model you use.

LLM Messages

A modern AI application commonly represents a conversation as messages.

For example:

System
User
Assistant
User
Assistant

In .NET, Microsoft.Extensions.AI provides ChatMessage and ChatRole for this purpose.

Example:

List<ChatMessage> messages =
[
    new ChatMessage(
        ChatRole.System,
        "You are a helpful C# programming assistant."),

    new ChatMessage(
        ChatRole.User,
        "What is dependency injection?")
];

System Messages

A system message describes the desired behavior of the assistant.

Example:

new ChatMessage(
    ChatRole.System,
    "You are a helpful .NET programming assistant.")

Conceptually:

System Instructions
       |
       v
User Request
       |
       v
LLM

System instructions are useful for consistent behavior, but they are not a substitute for application authorization or security controls.

User Messages

A user message represents the actual user input.

new ChatMessage(
    ChatRole.User,
    "Explain dependency injection.");

Typical flow:

System
   |
User
   |
Assistant
   |
User
   |
Assistant

Assistant Messages

The AI's previous responses can be retained as conversation history.

new ChatMessage(
    ChatRole.Assistant,
    "Dependency injection is...");

This allows the next request to include previous conversational context.

Conversation History

A simple chat application can use:

List<ChatMessage> chatHistory = [];

Then:

chatHistory.Add(
    new ChatMessage(
        ChatRole.User,
        userInput));

ChatResponse response =
    await chatClient.GetResponseAsync(
        chatHistory,
        cancellationToken: cancellationToken);

chatHistory.AddMessages(response);

The IChatClient API provides methods for complete and streaming responses, and Microsoft documents adding generated responses back into the message collection for multi-turn conversations.

Stateful vs Stateless LLM APIs

There are two broad approaches.

Stateless application-managed history

Your application sends the history each time:

Application
   |
   +---- Message 1
   +---- Message 2
   +---- Message 3
   +---- User Question
   |
   v
LLM

Provider-managed conversation state

Some APIs can maintain conversation identifiers or state.

IChatClient exposes ConversationId through response/options patterns for services that support stateful behavior. Microsoft's current documentation explains that applications can either retain history themselves or propagate a returned conversation identifier when the service maintains state.

For enterprise applications, explicitly deciding who owns conversation state is important.

LLM Context Construction

A .NET AI application often builds the final model input from several sources:

System Instructions
        +
Conversation History
        +
User Question
        +
RAG Context
        +
Tool Results
        +
Application Context
        |
        v
Final AI Request

This process is commonly called context construction.

Context Is Not the Same as Training

This distinction is extremely important.

Suppose you send:

Your company uses SQL Server.

as part of the prompt.

That does not mean the model has been retrained on your company.

Instead:

Your data
   |
   v
Current request context
   |
   v
LLM

The information affects the current generation through the provided context.

Context vs Fine-Tuning

Context:

Runtime information
      |
      v
Prompt
      |
      v
Model

Fine-tuning:

Training examples
      |
      v
Training process
      |
      v
Updated model

RAG:

External documents
      |
      v
Retrieval
      |
      v
Runtime context
      |
      v
Model

These are different techniques.

LLMs and RAG

RAG stands for Retrieval-Augmented Generation.

The basic architecture is:

Question
   |
   v
Retrieve relevant information
   |
   v
Add context to prompt
   |
   v
LLM
   |
   v
Answer

Later in this series, RAG will become a major topic.

At the LLM level, the key concept is that the model can generate an answer using context supplied by the application.

LLMs and Embeddings

Embeddings represent content as vectors.

For example:

"C# dependency injection"
       |
       v
Embedding Model
       |
       v
[0.12, -0.31, 0.67, ...]

These vectors are used for:

Semantic search
RAG
Similarity
Recommendations
Classification

The important distinction is:

LLM
    -> Generates / processes language

Embedding Model
    -> Converts content into vectors

The same AI application can use both.

LLMs and Tool Calling

LLMs do not normally have unrestricted access to your database or APIs.

Instead, your application can expose tools.

User
 |
 v
LLM
 |
 | Tool request
 v
.NET Tool
 |
 v
Database/API
 |
 v
Tool result
 |
 v
LLM
 |
 v
Answer

This turns an LLM from a pure text generator into a participant in an application workflow.

Microsoft's IChatClient supports tool calling through ChatOptions and function invocation middleware.

LLMs and Structured Output

A normal LLM response may look like:

The candidate has 5 years of experience
and works with C# and SQL Server.

An application may instead need:

{
  "experienceYears": 5,
  "skills": [
    "C#",
    "SQL Server"
  ]
}

Structured output allows generated information to fit a known schema.

IChatClient exposes typed response extensions for requesting output matching a C# type.

Example:

public sealed class CandidateProfile
{
    public int ExperienceYears { get; set; }

    public List<string> Skills { get; set; } = [];
}

Then conceptually:

CandidateProfile? profile =
    await chatClient.GetResponseAsync<CandidateProfile>(
        "Extract the candidate's experience and skills.");

The exact provider capabilities and structured-output constraints still need to be respected.

LLMs and Streaming

Normal generation:

Request
   |
   v
Wait
   |
   v
Complete response

Streaming:

Request
   |
   v
Token/chunk 1
   |
   v
Token/chunk 2
   |
   v
Token/chunk 3
   |
   v
...

IChatClient.GetStreamingResponseAsync returns an IAsyncEnumerable<ChatResponseUpdate> for incremental response delivery.

Example:

await foreach (
    ChatResponseUpdate update
    in chatClient.GetStreamingResponseAsync(
        "Explain dependency injection.",
        cancellationToken: cancellationToken))
{
    Console.Write(update.Text);
}

This is particularly useful for chat interfaces.

LLMs and Multimodal Input

Some modern LLM-capable systems support more than text.

Depending on the provider and model, inputs can include:

Text
Images
Audio
Other content types

IChatClient is designed to represent chat interactions with multimodal content such as text, images, and audio.

Conceptually:

               Chat Request
                     |
       +-------------+-------------+
       |             |             |
      Text         Image         Audio
       |             |             |
       +-------------+-------------+
                     |
                     v
                   Model

Later topics in this roadmap cover vision and speech in more detail.

LLMs and Model Providers

A .NET application can connect to many providers.

Microsoft's current .NET AI documentation lists integrations including:

OpenAI
Azure OpenAI
Azure AI Foundry
Ollama
Google Gemini
Amazon Bedrock

through compatible .NET AI abstractions.

Architecture:

                       IChatClient
                            |
         +------------------+------------------+
         |                  |                  |
       OpenAI             Azure              Ollama
                        OpenAI
         |                  |                  |
         +------------------+------------------+
                            |
                           LLM

Using LLMs with Microsoft.Extensions.AI

The simplest modern .NET abstraction is:

using Microsoft.Extensions.AI;

Then:

IChatClient chatClient = ...;

and:

ChatResponse response =
    await chatClient.GetResponseAsync(
        "What is an LLM?");

Microsoft's current quickstart uses this abstraction so that application code does not have to depend directly on a specific AI SDK.

Creating a .NET LLM Project

Create a console application:

dotnet new console -n DotNetLLM
cd DotNetLLM

Install the core AI abstractions:

dotnet add package Microsoft.Extensions.AI

For OpenAI:

dotnet add package OpenAI
dotnet add package Microsoft.Extensions.AI.OpenAI

The Microsoft quickstart currently uses Microsoft.Extensions.AI with OpenAI or Azure OpenAI integrations.

Configuring the OpenAI API Key

For local development, use an environment variable or User Secrets.

PowerShell:

$env:OPENAI_API_KEY = "YOUR_API_KEY"

Or initialize User Secrets:

dotnet user-secrets init

Then:

dotnet user-secrets set "AI:ApiKey" "YOUR_API_KEY"

Never place a production API key directly in source code.

The official OpenAI .NET library also recommends keeping credentials in secure configuration such as environment variables rather than source control.

Calling an LLM with OpenAI

Using the OpenAI SDK directly:

using OpenAI.Chat;

string apiKey =
    Environment.GetEnvironmentVariable("OPENAI_API_KEY")
    ?? throw new InvalidOperationException(
        "OPENAI_API_KEY is not configured.");

ChatClient client =
    new(
        model: "your-model-name",
        apiKey: apiKey);

ChatCompletion completion =
    await client.CompleteChatAsync(
        "Explain what an LLM is.");

Console.WriteLine(
    completion.Content[0].Text);

The official OpenAI .NET package currently provides a ChatClient and many other feature-specific clients.

Use a verified model identifier appropriate for your account and provider instead of copying a model name from an old tutorial.

Calling an LLM Through IChatClient

The provider-neutral approach is:

using Microsoft.Extensions.AI;
using OpenAI;

string apiKey =
    Environment.GetEnvironmentVariable("OPENAI_API_KEY")
    ?? throw new InvalidOperationException(
        "OPENAI_API_KEY is not configured.");

IChatClient chatClient =
    new OpenAIClient(apiKey)
        .GetChatClient("your-model-name")
        .AsIChatClient();

ChatResponse response =
    await chatClient.GetResponseAsync(
        "Explain what an LLM is.");

Console.WriteLine(response.Text);

The application now depends on:

IChatClient

rather than directly on:

OpenAI.Chat.ChatClient

This is the main architectural difference.

Creating a Simple LLM Service

Define:

public interface ILLMService
{
    Task<string> AskAsync(
        string prompt,
        CancellationToken cancellationToken = default);
}

Implementation:

using Microsoft.Extensions.AI;

public sealed class LLMService : ILLMService
{
    private readonly IChatClient _chatClient;

    public LLMService(IChatClient chatClient)
    {
        _chatClient = chatClient;
    }

    public async Task<string> AskAsync(
        string prompt,
        CancellationToken cancellationToken = default)
    {
        ChatResponse response =
            await _chatClient.GetResponseAsync(
                prompt,
                cancellationToken: cancellationToken);

        return response.Text;
    }
}

Now the application can depend on:

ILLMService

instead of any provider-specific class.

ASP.NET Core LLM Architecture

A simple Web API can use:

HTTP Request
     |
     v
LLMController
     |
     v
ILLMService
     |
     v
IChatClient
     |
     v
LLM Provider

Controller:

using Microsoft.AspNetCore.Mvc;

public sealed record LLMRequest(
    string Prompt);

[ApiController]
[Route("api/llm")]
public sealed class LLMController :
    ControllerBase
{
    private readonly ILLMService _llmService;

    public LLMController(
        ILLMService llmService)
    {
        _llmService = llmService;
    }

    [HttpPost("ask")]
    public async Task<IActionResult> Ask(
        LLMRequest request,
        CancellationToken cancellationToken)
    {
        if (string.IsNullOrWhiteSpace(
            request.Prompt))
        {
            return BadRequest(
                "Prompt is required.");
        }

        string response =
            await _llmService.AskAsync(
                request.Prompt,
                cancellationToken);

        return Ok(new
        {
            response
        });
    }
}

Configuring Dependency Injection

In Program.cs:

using Microsoft.Extensions.AI;
using OpenAI;

var builder =
    WebApplication.CreateBuilder(args);

string apiKey =
    builder.Configuration["AI:ApiKey"]
    ?? throw new InvalidOperationException(
        "AI:ApiKey is not configured.");

string model =
    builder.Configuration["AI:Model"]
    ?? throw new InvalidOperationException(
        "AI:Model is not configured.");

IChatClient chatClient =
    new OpenAIClient(apiKey)
        .GetChatClient(model)
        .AsIChatClient();

builder.Services.AddSingleton(chatClient);

builder.Services.AddScoped<
    ILLMService,
    LLMService>();

builder.Services.AddControllers();

var app =
    builder.Build();

app.MapControllers();

app.Run();

The official OpenAI .NET library documents its clients as thread-safe and suitable for singleton registration in ASP.NET Core dependency injection.

LLM Options

You can control model behavior using ChatOptions.

ChatOptions options = new()
{
    Temperature = 0.2f,
    MaxOutputTokens = 500
};

ChatResponse response =
    await chatClient.GetResponseAsync(
        "Explain async and await in C#.",
        options,
        cancellationToken);

Not every option is supported identically by every provider or model.

ChatOptions exposes common properties while allowing provider-specific options through additional properties or raw representations.

Model Selection

Model selection is an architectural concern.

Avoid:

new OpenAIClient(...)
    .GetChatClient("hard-coded-model-name");

in many different services.

Prefer:

Configuration
     |
     v
Model Selection
     |
     v
IChatClient

Example:

{
  "AI": {
    "Model": "your-model-name"
  }
}

Then:

string model =
    configuration["AI:Model"]
    ?? throw new InvalidOperationException(
        "AI model is not configured.");

This allows model changes without rebuilding business logic.

Choosing an LLM

Model selection should consider:

Quality requirements
Latency
Cost
Context requirements
Reasoning requirements
Structured output
Tool calling
Vision
Audio
Provider availability
Privacy requirements

Do not select a model based only on its name.

A model that is excellent for one workload may be unnecessary for another.

Small vs Large Models

A common architecture uses more than one model.

Simple Classification
        |
        v
Smaller / faster model

Complex reasoning
        |
        v
More capable model

This can improve:

Latency
Cost
Scalability

The correct choice should be validated with application-specific evaluations.

Model Routing

A more advanced architecture can route by task:

                Request
                   |
                   v
              AI Router
             /    |    \
            /     |     \
     Classification Chat  Reasoning
          |         |        |
        Model A   Model B   Model C

Microsoft.Extensions.AI also exposes routing functionality around IChatClient, including RoutingChatClient, for choosing among chat clients.

LLMs and Conversation Memory

An LLM does not automatically know what happened in previous application requests unless the required conversational state is available to the service.

The application may store:

Conversation ID
User ID
Messages
Summaries
Metadata

Architecture:

User
 |
 v
Conversation Service
 |
 v
Conversation Store
 |
 v
Context Builder
 |
 v
LLM

Conversation History Example

List<ChatMessage> history =
[
    new ChatMessage(
        ChatRole.System,
        "You are a helpful C# assistant."),

    new ChatMessage(
        ChatRole.User,
        "What is dependency injection?"),

    new ChatMessage(
        ChatRole.Assistant,
        "Dependency injection is..."),

    new ChatMessage(
        ChatRole.User,
        "Show me an example.")
];

ChatResponse response =
    await chatClient.GetResponseAsync(
        history);

This allows the model to use the earlier messages as current context.

LLMs and Context Reduction

Long histories need management.

A basic strategy is:

Old messages
     |
     v
Summarize
     |
     v
Summary + recent messages
     |
     v
LLM

For example:

Conversation:
100 messages

Context:
Summary of first 80
+
Recent 20

Microsoft's current Microsoft.Extensions.AI documentation describes experimental chat reducers including message-counting and summarization reducers.

LLMs and Prompt Engineering

Prompt engineering is the process of designing instructions and context so the model can perform the desired task.

A basic prompt:

Summarize this:
...

A more structured prompt:

Role:
You are a technical support assistant.

Task:
Summarize the issue.

Requirements:
- Be concise.
- Preserve technical facts.
- Do not invent information.

Input:
...

The model receives all of this as context.

Prompt vs Application Logic

Do not put business rules entirely inside prompts.

Bad architecture:

Prompt:
"You are allowed to refund only customers
who satisfy our refund policy..."

and then blindly trust the model.

Better:

Application
    |
    +---- Business Rules
    |
    +---- Authorization
    |
    +---- AI

The model can assist with interpretation, but deterministic business rules should stay in code.

LLMs and Grounding

An LLM can produce plausible output even when the required information is missing.

Grounding gives the model external context.

Question
 |
 v
Retrieve facts
 |
 v
Context
 |
 v
LLM
 |
 v
Grounded response

RAG is one of the primary grounding techniques.

LLMs and Hallucinations

A model can produce information that is incorrect or unsupported.

Therefore:

LLM Output
    |
    v
Validation
    |
    v
Business Decision

not:

LLM Output
    |
    v
Database Update

For important workflows, use:

Validation
Structured output
Grounding
Tool verification
Human approval

where appropriate.

LLMs and Deterministic Systems

LLMs are probabilistic systems.

Databases, validation rules, permissions, and arithmetic calculations can often be handled more deterministically.

A good architecture therefore uses the right component for the right job:

LLM
 -> Language understanding
 -> Generation
 -> Classification
 -> Summarization

Code
 -> Business rules
 -> Authorization
 -> Validation

Database
 -> Persistent state
 -> Exact retrieval

Search
 -> Retrieval

Tools
 -> External actions

LLMs and Tool Use

Suppose a user asks:

What is the status of order 10245?

The LLM may identify that it needs order information.

The application then allows:

GetOrderStatus(10245)

Architecture:

User
 |
 v
LLM
 |
 v
Tool Request
 |
 v
.NET Tool
 |
 v
Order Service
 |
 v
SQL Server
 |
 v
Tool Result
 |
 v
LLM
 |
 v
Final Answer

The model does not directly access SQL Server.

LLMs and SQL

A text-to-SQL system can follow:

User question
      |
      v
LLM
      |
      v
Generated SQL
      |
      v
SQL Validator
      |
      v
Authorization
      |
      v
Read-only database
      |
      v
Result
      |
      v
LLM explanation

Never treat generated SQL as automatically trusted.

LLMs and External APIs

Similarly:

LLM
 |
 v
Tool
 |
 v
External API
 |
 v
Response
 |
 v
LLM

Tools create a controlled bridge between probabilistic model behavior and deterministic application systems.

LLMs and Function Calling

Function calling means the model can request that a defined application function be invoked.

The conceptual flow is:

Messages
   |
   v
Model
   |
   +---- Normal response
   |
   +---- Tool request
              |
              v
         .NET function
              |
              v
          Tool result
              |
              v
             Model

Current IChatClient tooling supports this pattern.

LLMs and Streaming Architecture

For a web application:

Browser
   |
   v
ASP.NET Core
   |
   v
IChatClient
   |
   v
LLM

Streaming:

LLM
 |
 +---- Update 1
 |
 +---- Update 2
 |
 +---- Update 3
 |
 v
ASP.NET Core
 |
 v
Browser

This makes the user interface feel responsive even when the complete answer takes time to generate.

Streaming Example

await foreach (
    ChatResponseUpdate update
    in chatClient.GetStreamingResponseAsync(
        history,
        cancellationToken: cancellationToken))
{
    Console.Write(update.Text);
}

Microsoft recommends streaming through IAsyncEnumerable<ChatResponseUpdate> for incremental chat experiences.

LLMs and Local Models

Not all LLMs need to be accessed through a cloud provider.

Ollama provides a local model server, and Microsoft's current .NET quickstart shows using OllamaSharp with IChatClient to communicate with a local model.

Architecture:

.NET
 |
 v
OllamaSharp
 |
 v
Ollama
 |
 v
Local LLM

Local LLM Example

Install:

dotnet add package Microsoft.Extensions.AI
dotnet add package OllamaSharp

Then:

using Microsoft.Extensions.AI;
using OllamaSharp;

IChatClient chatClient =
    new OllamaApiClient(
        new Uri("http://localhost:11434/"),
        "your-local-model");

ChatResponse response =
    await chatClient.GetResponseAsync(
        "Explain dependency injection in C#.");

Console.WriteLine(response.Text);

Microsoft's current local-AI quickstart uses the same overall integration pattern.

Local LLM Advantages

Local LLMs can be useful for:

Local development
Offline scenarios
Private deployments
On-premises environments
Experimentation
Controlled data environments

But local models also require consideration of:

Hardware
Memory
Latency
Model quality
Deployment
Updates
Scaling

Cloud LLM Advantages

Cloud-hosted models can provide:

Managed infrastructure
Large models
Scalable capacity
Provider-managed APIs
Rapid model availability

But require attention to:

Network dependency
Cost
Data policies
Provider availability
Authentication
Rate limits

A hybrid architecture may use both.

Hybrid LLM Architecture

                  AI Router
                      |
            +---------+---------+
            |                   |
        Local LLM           Cloud LLM
            |                   |
      Private data        General workloads

Routing can depend on:

Data sensitivity
Cost
Quality
Latency
Availability
Large Language Models with .NET

LLMs and Azure OpenAI

Azure OpenAI can be integrated through the Azure SDK and adapted to IChatClient.

Conceptually:

.NET Application
      |
      v
AzureOpenAIClient
      |
      v
GetChatClient(...)
      |
      v
AsIChatClient()
      |
      v
IChatClient

Microsoft's current quickstart demonstrates this pattern with AzureOpenAIClient, DefaultAzureCredential, and AsIChatClient().

Azure OpenAI Example

using Azure.AI.OpenAI;
using Azure.Identity;
using Microsoft.Extensions.AI;

string endpoint =
    configuration["AzureOpenAI:Endpoint"]
    ?? throw new InvalidOperationException(
        "Azure OpenAI endpoint is not configured.");

string deployment =
    configuration["AzureOpenAI:Deployment"]
    ?? throw new InvalidOperationException(
        "Azure OpenAI deployment is not configured.");

IChatClient chatClient =
    new AzureOpenAIClient(
        new Uri(endpoint),
        new DefaultAzureCredential())
    .GetChatClient(deployment)
    .AsIChatClient();

This keeps the application-level interface the same.

LLM Abstraction Architecture

The application can therefore use:

                        IChatClient
                            |
           +----------------+----------------+
           |                |                |
         OpenAI          Azure OpenAI      Ollama
           |                |                |
       Cloud Model      Cloud Model       Local Model

The rest of the application can remain:

IAIService
RAGService
ConversationService
AgentService

LLMs and Semantic Kernel

Semantic Kernel can sit above basic model interaction.

Application
   |
   v
Semantic Kernel
   |
   v
AI Service
   |
   v
LLM

This becomes useful when the application needs:

Plugins
Functions
Prompt orchestration
Memory
Agent capabilities

The current Semantic Kernel documentation positions it as a higher-level framework rather than merely a raw LLM client.

LLMs and Agents

A basic LLM interaction is:

Question
   |
   v
LLM
   |
   v
Answer

An agent may be:

Goal
 |
 v
Agent
 |
 +---- LLM
 +---- Tool
 +---- RAG
 +---- Memory
 +---- Workflow
 |
 v
Result

This is why understanding basic LLM interaction should come before studying agents.

LLMs and MCP

MCP can connect an AI application to external tools and resources.

LLM Application
      |
      v
MCP Client
      |
      v
MCP Server
      |
 +----+------+------+
 |           |      |
API       Database  Files

The LLM is still the language/reasoning component.

MCP provides a standardized way for the application to access capabilities.

Microsoft's current .NET MCP documentation describes MCP as a client-server protocol for connecting AI applications with external tools and resources.

LLM Architecture in a Production Application

A realistic system may look like:

                             Client
                               |
                               v
                        ASP.NET Core API
                               |
                +--------------+--------------+
                |                             |
          Authentication                 Rate Limiting
                |                             |
                +--------------+--------------+
                               |
                               v
                         Application
                               |
          +--------------------+--------------------+
          |                    |                    |
          v                    v                    v
      AI Service            RAG Service          Tool Service
          |                    |                    |
          v                    v                    v
      IChatClient          Vector Store        Business APIs
          |
    +-----+-----+
    |     |     |
  Cache Telemetry Resilience
    |     |     |
    +-----+-----+
          |
          v
      Model Provider

This is the foundation for the remaining LLM topics.

LLMs and Error Handling

LLM API calls can fail because of:

Invalid credentials
Invalid model
Invalid request
Context limit
Rate limiting
Network failure
Timeout
Provider outage
Tool errors

A service layer should convert provider errors into appropriate application behavior.

Example:

try
{
    ChatResponse response =
        await _chatClient.GetResponseAsync(
            prompt,
            cancellationToken:
                cancellationToken);

    return response.Text;
}
catch (OperationCanceledException)
{
    throw;
}
catch (Exception ex)
{
    _logger.LogError(
        ex,
        "LLM request failed.");

    throw new AIServiceException(
        "The AI service could not process the request.",
        ex);
}

The client should not receive raw internal exceptions.

LLMs and Cancellation

LLM generation can take time.

Always propagate:

CancellationToken

through the layers.

HTTP Request
     |
     v
Controller
     |
     v
Application Service
     |
     v
IChatClient
     |
     v
Provider

The cancellation signal should flow through the entire chain.

LLMs and Timeouts

A production system should define appropriate timeouts.

Client timeout
Application timeout
Provider timeout
Tool timeout

Different AI workflows may require different policies.

A simple chat request might have a shorter timeout than a document-processing job.

LLMs and Retries

Transient failures may be retried.

Request
 |
 X
Retry
 |
 v
Provider
 |
 v
Success

Do not retry every failure.

Incorrect credentials or invalid requests should usually be handled as application errors rather than repeatedly retried.

LLMs and Rate Limits

Your application may have:

User rate limit
Tenant rate limit
Endpoint rate limit
Provider rate limit
Concurrent request limit

Architecture:

Client
 |
 v
Application Rate Limiter
 |
 v
LLM Service
 |
 v
Provider

This prevents a single user or tenant from consuming uncontrolled resources.

LLMs and Cost Control

A production application can track:

Model
Input tokens
Output tokens
Total tokens
Request count
Latency
Estimated cost

Architecture:

LLM Request
    |
    +---- Usage Tracking
    |
    +---- Budget Check
    |
    v
Provider

This becomes important for SaaS applications.

LLMs and Logging

Avoid logging full prompts automatically.

Prompts can contain:

Personal information
Private documents
Customer information
Business data
Source code
Credentials

Instead, log safe metadata:

_logger.LogInformation(
    "LLM request started. PromptLength={Length}",
    prompt.Length);

Possible safe telemetry:

Request ID
Model
Provider
Input length
Output length
Duration
Status
Token usage

LLMs and Observability

A production AI system should be observable.

                       LLM Request
                           |
            +--------------+--------------+
            |              |              |
          Logs          Metrics          Trace
            |              |              |
            +--------------+--------------+
                           |
                           v
                     AI Monitoring

Microsoft.Extensions.AI supports OpenTelemetry integration around chat clients.

LLMs and Determinism

An LLM is not a traditional deterministic function such as:

int Add(int x, int y)
{
    return x + y;
}

For:

Prompt -> LLM -> Answer

the exact output may vary based on:

Sampling
Model
Context
Provider implementation
Configuration

Therefore, tests should often evaluate properties of responses rather than compare every generated character exactly.

LLMs and Testing

Use different testing strategies.

Unit tests

Test:

Prompt building
Validation
Routing
Business rules
Authorization

using fake AI clients.

Integration tests

Test:

Real AI provider
Real SDK
Database
RAG
Tools

AI evaluation

Test:

Answer quality
Groundedness
Relevance
Completeness

This distinction becomes important later in the AI Testing section.

LLMs and Structured Application Data

A useful LLM architecture separates:

Natural language

from:

Application data

For example:

User Question
      |
      v
LLM
      |
      v
Structured Request
      |
      v
Business Service

Example:

{
  "customerId": 12345,
  "intent": "OrderStatus"
}

Then C# performs the actual operation.

LLMs and Business Rules

The LLM can identify:

Intent = OrderStatus

But the application determines:

Customer 12345 is authorized.

Then:

Business Service
      |
      v
SQL Server
      |
      v
Order Status

This keeps deterministic decisions inside application code.

LLMs and Application Architecture

A well-structured application can separate:

LLM Interaction
    |
    +---- Prompt
    +---- Context
    +---- Conversation
    +---- Tools
    +---- Output parsing

from:

Business Logic
    |
    +---- Rules
    +---- Permissions
    +---- Transactions
    +---- Data

This distinction becomes essential when LLM functionality grows.

LLMs and Background Processing

Some LLM workloads are better handled asynchronously:

Generate 10,000 summaries
Extract information from 5,000 documents
Classify millions of records
Generate embeddings

Architecture:

API
 |
 v
Queue
 |
 v
Worker
 |
 v
LLM

This avoids keeping an HTTP request alive for a long-running operation.

LLMs and Batch Processing

A batch workload can look like:

Records
   |
   v
Batch Scheduler
   |
   v
Worker Pool
   |
   v
LLM

Batch processing should account for:

Concurrency
Rate limits
Retries
Ordering
Idempotency
Cost

LLMs and Concurrency

One user may send multiple simultaneous requests.

A system can therefore require:

Concurrency limits
Queueing
Backpressure
Provider rate limits

Architecture:

Requests
   |
   v
Concurrency Controller
   |
   v
LLM Client

This is particularly important for high-volume APIs.

LLMs and Caching

Caching can reduce repeated model calls.

Question
 |
 v
Cache
 |
 +---- Hit ----> Answer
 |
 +---- Miss ---> LLM
                   |
                   v
                 Cache

However, caching must take context into account.

For example:

Same question
+
Different user

may require different answers.

LLMs and Semantic Caching

A semantic cache can compare vector representations of requests.

Question
   |
   v
Embedding
   |
   v
Similarity Search
   |
 +--+--+
 |     |
Match Miss
 |     |
 v     v
Cache  LLM
Answer

This becomes more useful in later RAG/vector topics.

LLMs and Privacy

LLM architecture should define:

What data can be sent?
Where is it processed?
How long is it retained?
Which provider receives it?
Can it be processed locally?

Possible architecture:

Public Data
    |
    v
Cloud LLM

Sensitive Data
    |
    v
Private / Local Processing

Privacy requirements should be determined by the application's domain and deployment environment.

LLMs and Security

Security should be enforced around the model.

User
 |
 v
Authentication
 |
 v
Authorization
 |
 v
Input Validation
 |
 v
AI
 |
 v
Output Validation

The prompt itself is not the security layer.

The application remains responsible for authorization and access control.

LLMs and Prompt Injection

When user input or retrieved documents are included in model context, malicious text can attempt to influence model behavior.

For example, retrieved content might contain instructions unrelated to the application's task.

Therefore:

User Input
   |
   v
Application Validation

Retrieved Data
   |
   v
Authorization

Tools
   |
   v
Tool Policy

The application should control capabilities independently of model-generated instructions.

LLMs and RAG Architecture

A complete RAG request might be:

Question
  |
  v
Query Embedding
  |
  v
Vector Search
  |
  v
Authorization Filter
  |
  v
Relevant Documents
  |
  v
Context Builder
  |
  v
LLM
  |
  v
Answer

This makes the LLM one component in a retrieval pipeline.

LLMs and AI Agents

Once tools, memory, RAG, planning, and workflows are combined:

               Agent
                 |
     +-----------+-----------+
     |           |           |
    LLM         RAG        Tools
                              |
                              v
                       External Systems

The LLM is still the central language/reasoning engine.

The surrounding software provides the capabilities.

LLMs and MCP

With MCP:

Agent
 |
 v
MCP Client
 |
 v
MCP Server
 |
 +---- Search
 +---- Database
 +---- Files
 +---- Business APIs

This creates a standardized capability layer around the model.

The Complete LLM Application Stack

A mature application can therefore look like:

                     User
                      |
                      v
                Client Application
                      |
                      v
                ASP.NET Core
                      |
                      v
               Application Layer
                      |
       +--------------+--------------+
       |              |              |
       v              v              v
    AI Service      RAG Service    Tool Service
       |              |              |
       v              v              v
    IChatClient    Vector Store   Business APIs
       |              |
       |          Embeddings
       |              |
       +------+-------+
              |
        AI Middleware
              |
      +-------+-------+
      |       |       |
    Cache Telemetry Resilience
              |
              v
          LLM Provider

LLMs and Model Abstraction

One of the most useful .NET architectural ideas is to keep:

Application

independent from:

Provider SDK

For example:

Application
     |
     v
IChatClient
     |
     +---- OpenAI
     +---- Azure OpenAI
     +---- Ollama

The official OpenAI .NET SDK supports direct feature clients such as ChatClient, ResponsesClient, EmbeddingClient, AudioClient, and VectorStoreClient, while Microsoft.Extensions.AI provides common abstractions over provider implementations.

LLMs and the OpenAI Responses API

The official OpenAI .NET library currently exposes a ResponsesClient in addition to ChatClient. The SDK documentation describes ResponsesClient under OpenAI.Responses.

This illustrates an important point:

LLM application development

is not necessarily limited to one API style.

Providers may expose:

Chat APIs
Responses APIs
Embeddings APIs
Audio APIs
Image APIs
Realtime APIs

The correct API depends on the application requirement.

Chat API vs General Model API

A chat-oriented abstraction represents interactions as messages:

System
User
Assistant

A broader response-oriented API may support richer capabilities.

The application architecture can therefore be:

Application
   |
   +---- Common AI abstraction
   |
   +---- Provider-specific client when needed

This is one reason not every advanced capability should be forced through one interface.

Direct SDK vs IChatClient

There are two valid approaches.

Direct provider SDK

Application
   |
   v
OpenAI ChatClient
   |
   v
OpenAI

Advantages:

Direct provider control
Provider-specific features
Latest provider functionality

IChatClient

Application
   |
   v
IChatClient
   |
   v
Provider Client

Advantages:

Provider abstraction
Testability
Dependency injection
Common middleware
Easier provider switching

The appropriate choice depends on the architecture.

LLMs and ASP.NET Core Dependency Injection

The OpenAI .NET library documents its clients as thread-safe and suitable for singleton registration in ASP.NET Core.

For abstraction-based applications:

builder.Services.AddSingleton(
    chatClient);

Then:

public AIService(
    IChatClient chatClient)
{
    _chatClient = chatClient;
}

This keeps client creation out of business services.

LLMs and Configuration

Example configuration:

{
  "AI": {
    "Provider": "OpenAI",
    "Model": "your-model-name"
  }
}

For Azure:

{
  "AI": {
    "Provider": "AzureOpenAI",
    "Endpoint": "https://your-resource.openai.azure.com",
    "Deployment": "your-deployment-name"
  }
}

Keep credentials outside source-controlled configuration.

LLMs and Environment Variables

For local development:

$env:OPENAI_API_KEY = "YOUR_API_KEY"

For production:

Environment variables
Managed identity
Secret store
Key Vault

The official OpenAI SDK recommends secure credential storage rather than placing secrets in source control.

LLMs and User Experience

LLM architecture also affects UI behavior.

Complete response

User -> Wait -> Full answer

Streaming

User -> Answer begins immediately
       |
       +---- More text
       |
       +---- More text

For chat applications, streaming often creates a more responsive experience.

Microsoft's current IChatClient guidance explicitly describes streaming as a natural fit for AI user experiences.

LLMs and Response Validation

A model response can be:

Syntactically valid
but
Semantically incorrect

Therefore:

Model Response
      |
      v
Deserialization
      |
      v
Validation
      |
      v
Business Rules
      |
      v
Application Action

Structured output reduces parsing complexity but does not eliminate application validation.

LLMs and AI Quality

A successful API call does not necessarily mean successful AI behavior.

For example:

HTTP 200
JSON valid
Response generated

while:

Answer incorrect

Therefore AI systems need:

Evaluation
Test datasets
Regression testing
Grounding checks

LLMs and Evaluation

A useful evaluation architecture is:

Question Dataset
       |
       v
LLM Application
       |
       v
Generated Answers
       |
       v
Evaluation
       |
       v
Quality Report

This will become a dedicated topic later in the roadmap.

LLMs and Enterprise Architecture

A larger enterprise architecture may look like:

                             Users
                               |
                               v
                         API Gateway
                               |
                               v
                       ASP.NET Core
                               |
          +--------------------+--------------------+
          |                    |                    |
          v                    v                    v
      AI Service           RAG Service          Agent Service
          |                    |                    |
          v                    v                    v
     IChatClient          Vector Store          Tools/MCP
          |                    |
          +----------+---------+
                     |
                     v
                   LLM

Supporting systems:

SQL Server
Redis
Blob Storage
Message Broker
Vector Database
Secret Management
OpenTelemetry
Evaluation

LLMs and the Cloud

A cloud-hosted architecture:

ASP.NET Core
     |
     v
Cloud AI Endpoint
     |
     v
Hosted Model

Benefits can include:

Managed infrastructure
Scalability
Large model availability
Operational tooling

LLMs and On-Premises

An on-premises architecture:

ASP.NET Core
     |
     v
Local Model Server
     |
     v
Local LLM

Potential advantages:

Data control
Network isolation
Local processing

Potential challenges:

Hardware
Capacity
Model management
Scaling
Operations

LLMs and Hybrid Cloud

A hybrid architecture can route workloads:

Application
    |
    v
AI Router
   / \
  /   \
Local Cloud

This enables:

Private data -> Local/private model
General data -> Cloud model

but introduces additional routing and operational complexity.

LLM Architecture Decision Checklist

Before integrating an LLM, answer:

What model capability is needed?
Which provider should host it?
Does the application require provider independence?
How large can the context become?
How will conversation history be stored?
Will RAG be required?
Will tools be required?
Will structured output be required?
Will responses be streamed?
What happens when the provider fails?
How will tokens be tracked?
How will cost be controlled?
How will AI quality be tested?
What data can leave the application?

These questions should be answered before designing the production integration.

Practical Project: .NET LLM Chat Application

Create:

dotnet new webapi -n DotNetLLMChat
cd DotNetLLMChat

Add:

dotnet add package Microsoft.Extensions.AI
dotnet add package Microsoft.Extensions.AI.OpenAI
dotnet add package OpenAI

Configuration:

{
  "AI": {
    "Model": "your-model-name"
  }
}

User Secret:

dotnet user-secrets init
dotnet user-secrets set "AI:ApiKey" "YOUR_API_KEY"

Service:

using Microsoft.Extensions.AI;

public interface ILLMService
{
    Task<string> AskAsync(
        string prompt,
        CancellationToken cancellationToken = default);
}

Implementation:

using Microsoft.Extensions.AI;

public sealed class LLMService :
    ILLMService
{
    private readonly IChatClient _chatClient;

    public LLMService(
        IChatClient chatClient)
    {
        _chatClient = chatClient;
    }

    public async Task<string> AskAsync(
        string prompt,
        CancellationToken cancellationToken = default)
    {
        ChatResponse response =
            await _chatClient.GetResponseAsync(
                prompt,
                cancellationToken: cancellationToken);

        return response.Text;
    }
}

Registration:

using Microsoft.Extensions.AI;
using OpenAI;

var builder =
    WebApplication.CreateBuilder(args);

string apiKey =
    builder.Configuration["AI:ApiKey"]
    ?? throw new InvalidOperationException(
        "AI:ApiKey is not configured.");

string model =
    builder.Configuration["AI:Model"]
    ?? throw new InvalidOperationException(
        "AI:Model is not configured.");

IChatClient chatClient =
    new OpenAIClient(apiKey)
        .GetChatClient(model)
        .AsIChatClient();

builder.Services.AddSingleton(
    chatClient);

builder.Services.AddScoped<
    ILLMService,
    LLMService>();

builder.Services.AddControllers();

var app =
    builder.Build();

app.MapControllers();

app.Run();

Controller:

using Microsoft.AspNetCore.Mvc;

public sealed record AskRequest(
    string Prompt);

[ApiController]
[Route("api/llm")]
public sealed class LLMController :
    ControllerBase
{
    private readonly ILLMService _llmService;

    public LLMController(
        ILLMService llmService)
    {
        _llmService = llmService;
    }

    [HttpPost("ask")]
    public async Task<IActionResult> Ask(
        AskRequest request,
        CancellationToken cancellationToken)
    {
        if (string.IsNullOrWhiteSpace(
            request.Prompt))
        {
            return BadRequest(
                "Prompt is required.");
        }

        string response =
            await _llmService.AskAsync(
                request.Prompt,
                cancellationToken);

        return Ok(new
        {
            response
        });
    }
}

Request:

POST /api/llm/ask
Content-Type: application/json
{
  "prompt": "Explain dependency injection in ASP.NET Core."
}

Architecture:

Client
  |
  v
ASP.NET Core
  |
  v
ILLMService
  |
  v
IChatClient
  |
  v
OpenAI
  |
  v
LLM

Practical Project: Conversation-Based LLM Application

Extend the project with:

Conversation
ConversationMessage
Conversation repository
context manager

Architecture:

User
 |
 v
Conversation API
 |
 +---- Load History
 |
 +---- Add User Message
 |
 +---- Context Management
 |
 v
IChatClient
 |
 v
LLM
 |
 v
Store Assistant Response

This becomes the foundation for the next roadmap topics:

Chat Models with .NET
AI Chat Applications with .NET
AI Conversation History with .NET
AI Context Management with .NET

Practical Project: Local LLM Version

Replace the OpenAI provider with Ollama.

Install:

dotnet add package OllamaSharp

Configure:

using Microsoft.Extensions.AI;
using OllamaSharp;

IChatClient chatClient =
    new OllamaApiClient(
        new Uri("http://localhost:11434/"),
        "your-local-model");

The application service remains:

ILLMService
    |
    v
IChatClient

Only infrastructure changes.

Microsoft's current local-model quickstart demonstrates this approach with OllamaSharp, IChatClient, conversation history, and streaming.

Practical Project: Multi-Provider LLM Application

Support:

OpenAI
Azure OpenAI
Ollama

Architecture:

                      ILLMService
                           |
                       IChatClient
                           |
           +---------------+---------------+
           |               |               |
         OpenAI          Azure           Ollama
                        OpenAI

Configuration:

{
  "AI": {
    "Provider": "OpenAI",
    "Model": "your-model-name"
  }
}

The application logic remains provider-neutral.

Practical Project: LLM + RAG

Build:

PDF
 |
 v
Document Ingestion
 |
 v
Chunks
 |
 v
Embeddings
 |
 v
Vector Database

Then:

Question
 |
 v
Vector Search
 |
 v
Context
 |
 v
IChatClient
 |
 v
LLM
 |
 v
Answer

This will become the detailed RAG implementation later in the roadmap.

Practical Project: LLM + Tool Calling

Create:

GetWeather
GetOrder
GetCustomer
SearchKnowledgeBase

Architecture:

User
 |
 v
LLM
 |
 +---- GetOrder
 |
 +---- SearchKnowledgeBase
 |
 v
Answer

The .NET application remains responsible for:

Authorization
Validation
Execution
Auditing

Practical Project: LLM + Structured Output

Create:

public sealed class InterviewQuestion
{
    public string Question { get; set; } =
        string.Empty;

    public string Difficulty { get; set; } =
        string.Empty;

    public List<string> Skills { get; set; } = [];
}

Then request structured output from an IChatClient.

The architecture:

Prompt
 |
 v
LLM
 |
 v
Structured Output
 |
 v
C# Type
 |
 v
Validation

This is useful for:

Resume processing
Document extraction
Classification
Form generation
API automation

Practical Project: LLM-Powered API

A production API might provide:

POST /api/llm/chat
POST /api/llm/summarize
POST /api/llm/classify
POST /api/llm/extract
POST /api/llm/stream

Architecture:

                    LLM API
                       |
       +---------------+---------------+
       |               |               |
      Chat         Summarize        Extract
       |               |               |
       +---------------+---------------+
                       |
                  IChatClient
                       |
                    Provider

This turns raw LLM access into reusable application functionality.

Common Mistakes with LLMs in .NET

Treating the LLM as a database

An LLM should not replace authoritative application data.

Use:

Database -> facts
LLM -> language generation

Putting secrets in source code

Never:

var key = "real-secret";

Use secure configuration.

Sending unlimited conversation history

Use:

Summarization
Trimming
Retrieval
Context reduction

Assuming model output is always correct

Use:

Validation
Grounding
Structured output
Tool verification
Evaluation

Putting business authorization in prompts

Authorization belongs in application code.

Hard-coding one provider everywhere

Use abstractions where provider independence is valuable.

Selecting models based only on popularity

Evaluate the model against:

Quality
Latency
Cost
Required capabilities
Context requirements

Ignoring streaming

For interactive chat, streaming can significantly improve perceived responsiveness.

Treating temperature as a quality setting

Temperature changes sampling behavior; it does not guarantee better or worse answers.

Ignoring context size

Large prompts can increase:

Latency
Usage
Cost
Context pressure

LLM Best Practices

Keep prompts explicit

State:

Role
Task
Constraints
Expected output
Relevant context

Keep business rules in code

Use the model for language and reasoning tasks.

Use application logic for deterministic business rules.

Keep provider code isolated

Use:

IChatClient
IAIService

where useful.

Use streaming for interactive applications

Especially for:

Chat
Assistants
Long responses

Use structured output for machine-readable data

Do not depend on arbitrary prose parsing when a schema is available.

Manage context intentionally

Use:

History trimming
Summaries
RAG
Context selection

Monitor AI usage

Track:

Tokens
Latency
Model
Provider
Failures
Tool calls

Evaluate model behavior

Use datasets and AI evaluations rather than relying only on unit tests.

Frequently Asked Questions

What is an LLM?

An LLM is a machine-learning model trained on large amounts of data to learn patterns in language and generate sequences of tokens.

What does an LLM actually generate?

An LLM generates tokens sequentially, with each token selected based on the current context and model probabilities.

What is a token?

A token is a unit used by the model's tokenizer to represent text. A token may be a word, word fragment, punctuation, or another character sequence.

Why are tokens important?

They affect:

Context limits
Usage
Cost
Latency

What is a context window?

The maximum amount of token context the model can process for a request, including relevant input/context and generated output.

Does conversation history automatically exist inside the model?

Not necessarily.

The application or provider must make the necessary conversation state available.

What is the difference between an LLM and an embedding model?

An LLM primarily processes and generates language.

An embedding model converts content into vector representations useful for semantic similarity and search.

What is RAG?

Retrieval-Augmented Generation combines retrieval of external information with LLM generation.

What is tool calling?

It allows a model to request that an application-defined function be executed and then use the returned result.

What is structured output?

It is a response format designed to conform to an application-defined schema or C# type.

What is streaming?

Streaming sends parts of the model response incrementally instead of waiting for the complete response.

Does IChatClient support streaming?

Yes. IChatClient provides GetStreamingResponseAsync, which returns asynchronous response updates.

Can IChatClient work with local LLMs?

Yes. Microsoft's current Ollama quickstart shows OllamaSharp implementing IChatClient.

Can IChatClient work with OpenAI?

Yes. Microsoft documents adapting the OpenAI .NET client to IChatClient.

Can IChatClient work with Azure OpenAI?

Yes. Microsoft's current quickstart shows AzureOpenAIClient(...).GetChatClient(...).AsIChatClient().

What is the difference between ChatClient and IChatClient?

ChatClient from the OpenAI SDK is provider-specific.

IChatClient is the provider-neutral abstraction from Microsoft.Extensions.AI.

Should every application use IChatClient?

No.

A provider-specific application can use its SDK directly.

IChatClient becomes especially useful when you want common abstractions, dependency injection, middleware, testing, or provider flexibility.

What is temperature?

A sampling parameter that can influence variation in generated output. Its exact effect depends on the model and provider.

Can I use more than one LLM?

Yes.

An application can route different workloads to different models or providers.

Can an LLM access my database directly?

It should not normally have unrestricted direct access.

Use controlled tools and application services.

Can an LLM enforce authorization?

No.

Authorization should be enforced by application code.

Can an LLM replace SQL Server?

No.

The LLM and database have fundamentally different responsibilities.

Why can an LLM produce incorrect answers?

Because generated output is based on learned patterns and supplied context rather than guaranteed retrieval of ground truth.

How can incorrect answers be reduced?

Use:

Grounding
RAG
Tools
Structured output
Validation
Evaluation

depending on the use case.

Interview Questions

What is a Large Language Model?

A machine-learning model designed to process and generate language by learning statistical relationships among token sequences.

What is autoregressive generation?

A generation process where the model predicts the next token based on the tokens already present in the sequence.

What is tokenization?

The conversion of text into the token units consumed by the model.

What is a context window?

The maximum token context that can be processed by a model for a request.

Why does context management matter?

Large conversation histories and retrieved documents consume context and can increase usage and latency.

What is the role of IChatClient?

It provides a common .NET abstraction for interacting with chat-capable AI services.

What is the role of ChatMessage?

It represents a message within an AI conversation and can carry roles such as system, user, and assistant.

What is the difference between training and inference?

Training learns model parameters from data.

Inference uses the trained model to generate an output for an input.

What is RAG?

A technique that retrieves external information and includes it as context for LLM generation.

What is tool calling?

A mechanism where the model requests execution of an application-defined function.

What is structured output?

An approach where the model response is constrained or requested to match a defined schema or type.

What is streaming?

The incremental delivery of model-generated output while generation is still occurring.

Why should API clients be registered through DI?

To centralize configuration, reuse clients, improve testability, and keep client creation outside business logic.

Why are provider abstractions useful?

They reduce coupling between the application and a particular AI provider.

What is an LLM hallucination?

An AI-generated statement that is incorrect, unsupported, or fabricated.

How does RAG reduce unsupported answers?

By retrieving relevant external information and supplying it as context for generation.

Why can prompt size affect cost?

Because the input context is represented as tokens and can contribute to model usage.

Why can large context windows still require context management?

Because larger windows still have finite limits and larger inputs can increase latency, usage, and irrelevant context.

Why should AI output be validated?

Because a syntactically valid response can still violate business or application requirements.

What is model routing?

Selecting different AI models or providers based on workload characteristics or application policy.

Exercises

Exercise 1: Basic LLM Application

Create a .NET console application that:

Reads user input
Calls an LLM
Prints the response

Exercise 2: Chat History

Maintain:

List<ChatMessage>

and create a multi-turn conversation.

Exercise 3: System Instructions

Add:

System
User
Assistant

messages and observe how the model's behavior changes.

Exercise 4: ChatOptions

Experiment with:

Temperature
Maximum output tokens
Model selection

Use only settings supported by the selected provider/model.

Exercise 5: Streaming

Replace:

GetResponseAsync(...)

with:

GetStreamingResponseAsync(...)

and display the response as it arrives.

Exercise 6: ASP.NET Core LLM API

Create:

POST /api/llm/ask

with:

{
  "prompt": "Explain ASP.NET Core dependency injection."
}

Exercise 7: Provider Switching

Support:

OpenAI
Ollama

with the same:

ILLMService
IChatClient

architecture.

Exercise 8: Context Management

Store the last 10 messages.

Then add a summary for older messages.

Architecture:

Recent Messages
+
Conversation Summary

Exercise 9: Structured Output

Create:

public sealed class ProductAnalysis
{
    public string Category { get; set; } =
        string.Empty;

    public List<string> Features { get; set; } = [];
}

Ask the model to produce the structured result and validate it.

Exercise 10: Tool Calling

Create:

GetProductDetails

as a tool.

Allow the model to request the function and then use the result.

Exercise 11: Local LLM

Run a local model through Ollama and connect to it using OllamaSharp.

Exercise 12: Multi-Model Router

Create:

Simple request -> Model A
Complex request -> Model B

behind one application service.

Exercise 13: LLM Evaluation

Create 20 test prompts.

Record:

Question
Expected characteristics
Model output
Evaluation result

Compare two models against the same dataset.

Practical Project: .NET LLM Assistant

Build an assistant with:

ASP.NET Core
SQL Server
Redis
IChatClient
OpenAI / Azure OpenAI

Features:

Login
Conversation history
Streaming
Context management
Usage tracking

Architecture:

                        User
                         |
                         v
                   ASP.NET Core
                         |
               +---------+---------+
               |                   |
         Authentication       Conversation
                                  |
                                  v
                           Context Manager
                                  |
                                  v
                            IChatClient
                                  |
                      +-----------+-----------+
                      |           |           |
                    Cache     Telemetry    Resilience
                      |           |           |
                      +-----------+-----------+
                                  |
                                  v
                              LLM Provider

Practical Project: Local + Cloud LLM Assistant

Add:

Ollama

alongside:

OpenAI

Architecture:

                     AI Router
                         |
              +----------+----------+
              |                     |
         Local LLM               Cloud LLM
              |                     |
           Ollama                 OpenAI

Policies can determine which requests use each provider.

Practical Project: Enterprise LLM Platform

A larger system:

                               Client
                                 |
                                 v
                           API Gateway
                                 |
                                 v
                          ASP.NET Core
                                 |
               +-----------------+-----------------+
               |                                   |
          Conversation                          AI Service
               |                                   |
               v                                   v
           SQL Server                         IChatClient
                                                   |
                         +-------------------------+-------------------------+
                         |                         |                         |
                       Cache                   Telemetry                Resilience
                         |                         |                         |
                         +-------------------------+-------------------------+
                                                   |
                                              Model Router
                                                   |
                              +--------------------+--------------------+
                              |                    |                    |
                            OpenAI              Azure                Ollama
                                               OpenAI

Later this can evolve into:

RAG
Tools
Agents
MCP
Evaluation
Distributed Workers

LLM Learning Path

This article introduced the fundamentals.

The next topics will go deeper into individual programming techniques:

#17 Using LLMs in C#
#18 Calling LLM APIs from .NET
#19 Chat Models with .NET
#20 Text Generation with .NET
#21 AI Text Generation with C#
#22 AI Chat Applications with .NET
#23 Streaming AI Responses with .NET
#24 AI Conversation History with .NET
#25 AI Context Management with .NET
#26 AI Prompt Management with .NET
#27 System Prompts with .NET
#28 Prompt Templates with .NET
#29 Structured AI Responses with .NET
#30 JSON Responses from AI Models with C#

After these topics, the course moves into:

OpenAI
Azure OpenAI
Microsoft AI
Semantic Kernel
RAG
Embeddings
Vector Databases
Local AI
Vision
Speech
Documents
Databases
Agents
MCP
Blazor
.NET MAUI
Automation
Security
Testing
Production

Key Takeaways

A Large Language Model is fundamentally a model that processes token sequences and generates new token sequences based on learned patterns.

The simplified flow is:

Text
 |
 v
Tokens
 |
 v
LLM
 |
 v
Next-token prediction
 |
 v
Output tokens
 |
 v
Text

A .NET application's job is not to implement the model itself.

It provides:

Input
Context
Conversation
Tools
Data
Security
Validation
Infrastructure

around the model.

The modern .NET abstraction is:

Application
     |
     v
IChatClient
     |
     v
LLM Provider
     |
     v
Model

The provider can be changed:

OpenAI
Azure OpenAI
Ollama
Other supported providers

while application code can remain largely unchanged. Microsoft's current .NET AI documentation explicitly positions Microsoft.Extensions.AI as the common abstraction for this style of model integration.

Conclusion

Large Language Models are the core technology behind many modern AI applications, but an LLM by itself is not an application.

The model provides capabilities such as:

Language understanding
Text generation
Summarization
Classification
Code generation
Reasoning
Structured generation

Your .NET application provides everything around the model:

Authentication
Authorization
Conversation state
Context
RAG
Tools
Databases
Caching
Resilience
Observability
Evaluation

The architecture can therefore be viewed as:

                         .NET Application
                                |
                    +-----------+-----------+
                    |                       |
              Business Logic            AI Layer
                                            |
                                            v
                                       IChatClient
                                            |
                         +------------------+------------------+
                         |                  |                  |
                       OpenAI             Azure              Ollama
                                         OpenAI
                         |                  |                  |
                         +------------------+------------------+
                                            |
                                           LLM

The most important concepts to remember are:

Tokens
Context windows
Inference
Messages
Conversation history
Prompt construction
Sampling
Streaming
Structured output
Embeddings
RAG
Tool calling
Provider abstraction

A strong .NET LLM application does not simply call a model.

It carefully manages the context given to the model, validates the output it receives, controls which actions the model can request, protects user data, handles provider failures, tracks usage, and evaluates the quality of generated results.

That foundation is important because the next several topics move from the theory of LLMs into practical C# programming.

The next topic is:

Using LLMs in C#

where the focus shifts from understanding LLMs to writing C# code that directly works with them.

Post a Comment

Previous Post Next Post