Latest Posts
Microsoft AI-103: Prompt Injection Defenses on Azure
Prompt injection is the attempt to make an AI system follow instructions that conflict with the developer’s intended behavior. The attack can come directly from the user or indirectly from content the system reads, such as documents, websites, emails, search results, or tool output. Indirect injection is especially important for RAG and agents because untrusted content can enter the model as if it were evidence. Azure AI Content Safety Prompt Shields detects user prompt attacks and document attacks. Detection is useful, but it is one layer in a broader defense….
Microsoft AI-103: Private Networking for Azure AI
Private networking for Azure AI is about removing unnecessary public exposure without confusing network isolation with identity or authorization. A private endpoint changes how traffic reaches a service. It does not decide who is allowed to call the service, which model a workload may use, or which data an agent can retrieve. Production architecture needs both network controls and identity controls. Microsoft Foundry can be configured with public network access disabled and private endpoints for inbound access. Supporting services such as Azure AI Search, Storage, Key Vault, Cosmos DB, and…
Microsoft AI-103: Online Evaluation for AI Systems
Online evaluation closes the gap between a controlled benchmark and the behavior users actually experience. Prelaunch datasets are essential for release gates, but production traffic introduces new phrasing, unexpected tool paths, changing knowledge, rare edge cases, and interactions the development team never wrote down. A mature AI system therefore needs a way to measure deployed behavior without treating every quality issue as a support anecdote. Microsoft Foundry can evaluate deployed interactions and Application Insights traces with the same evaluation framework used for offline quality analysis. Some production-trace and conversation evaluation…
Microsoft AI-103: Monitoring Model Drift in Azure ML
Model drift monitoring answers a production question that offline accuracy cannot: has the relationship between the model, its inputs, and the environment changed enough that yesterday’s validation no longer describes today’s behavior? Azure Machine Learning model monitoring helps teams compare production data with a reference and track signals such as data drift, prediction drift, data quality, feature attribution drift, and model performance where ground truth is available. Drift is not automatically a failure. A seasonal change, new customer segment, or product launch can move the input distribution while the model…
Microsoft AI-103: Managing Agent Memory on Azure
Agent memory is useful when an AI system needs continuity beyond one request or one session. It can remember a user’s stable preferences, summarize prior interactions, or retain reusable procedures learned across conversations. That same persistence creates governance questions: what deserves to be remembered, how long should it live, who can read it, and how can it be corrected or deleted? Microsoft Foundry Agent Service currently provides managed long-term memory in preview. The service distinguishes short-term conversational context from persistent memory and supports user profile memory, chat summary memory, and…
Microsoft AI-103: MLOps and GenAIOps Together
MLOps and GenAIOps are not competing operating models. They solve overlapping parts of the AI lifecycle. MLOps grew around training, registering, deploying, and monitoring predictive models. GenAIOps extends those practices to systems where behavior also depends on foundation-model selection, prompts, retrieval, agent orchestration, safety controls, and evaluation. Microsoft’s Azure Well-Architected guidance explicitly treats GenAIOps as a specialization that complements established DevOps, DataOps, and MLOps practices. The operational stages remain familiar: prepare data, validate behavior, automate release, observe production, and maintain the system. The assets and quality checks expand. That makes…
Microsoft AI-103: Latency Tuning for Azure AI Apps
Latency in an Azure AI application is the sum of several systems, not a single model response time. A user can wait on authentication, retrieval, prompt assembly, model queueing, time to first token, token generation, tool calls, safety checks, and network hops. Tuning only the model endpoint can leave most of the delay untouched. Microsoft’s Azure OpenAI guidance separates system throughput from per-call latency and emphasizes that output token count is usually one of the strongest latency drivers. Streaming can improve perceived responsiveness even when the total generation time is…
Microsoft AI-103: Hybrid Search in Azure AI Search
Hybrid search is one of the most practical retrieval patterns in Azure AI Search because enterprise questions rarely fit neatly into either keyword search or vector search. Users mix exact identifiers with natural language. They search for product codes, policy names, dates, abbreviations, error messages, and concepts that may be phrased differently in the source material. A retrieval system that commits to only one matching strategy gives up useful evidence before ranking even begins. Azure AI Search runs full-text and vector queries in parallel inside a single hybrid request and…
Microsoft AI-103: Handling Hallucinations in Azure AI
Hallucination is a useful shorthand for a generated statement that is unsupported, fabricated, or inconsistent with the evidence the application should have used. It is not one failure with one fix. A bad answer can begin in retrieval, prompt construction, model behavior, tool output, stale data, or the application’s decision to answer when it should abstain. Azure provides controls at several layers. RAG can ground generation in enterprise evidence. Microsoft Foundry includes evaluators such as groundedness, relevance, response completeness, and task completion. Foundry tracing can connect model output to retrieval,…
Microsoft AI-103: Grounding Azure AI with Enterprise Data
Enterprise grounding is not simply connecting a model to documents. A production system has to find the right evidence, respect the caller’s permissions, preserve source attribution, handle changing data, and decide whether information should be indexed or queried live. The model only sees the result of that data architecture. Azure AI Search now underpins Foundry IQ, Microsoft’s managed knowledge layer for reusable, permission-aware knowledge bases in Microsoft Foundry. Azure AI Search also supports agentic retrieval with knowledge sources that can connect indexed search content and, in some cases, remote systems….
Microsoft AI-103: GenAIOps on Azure
GenAIOps is the operating discipline for generative AI systems after the prototype stage. It combines versioning, evaluation, tracing, monitoring, release control, cost management, and feedback so that model-driven behavior can change without becoming unpredictable. The focus is not one Azure product. It is the delivery loop that connects development evidence to production evidence. Microsoft Foundry now supports reusable evaluations, agent tracing with OpenTelemetry, Application Insights integration, monitoring dashboards, continuous evaluation, and analysis of deployed interactions. Azure DevOps or GitHub-based delivery pipelines can then use those signals to control promotion. The…
Microsoft AI-103: From AI Prototype to Production on Azure
An AI prototype proves that a capability is possible. Production proves that the capability can be trusted repeatedly under real users, real data, real load, changing models, operational failures, and governance requirements. The gap between those two states is where many AI projects become expensive: the demo worked, but the system was never designed to be operated. Microsoft Foundry now brings model deployment, agents, evaluation, tracing, monitoring, and governance closer together, while Azure services provide identity, networking, queues, search, storage, observability, and cost management. The production task is to turn…
Microsoft AI-103: Event-Driven AI Workflows on Azure
Event-driven architecture is a natural fit for AI work that begins when something changes rather than when a user waits on a synchronous request. A document arrives, a ticket changes state, a message is published, a file is written, or a business event crosses a threshold. Those events can trigger enrichment, classification, evaluation, retrieval updates, agent work, or downstream automation without holding the original caller open. Azure provides several services for different parts of that pattern. Event Grid routes events from sources to handlers. Service Bus queues and topics provide…
Microsoft AI-103: Durable AI Workflows with Queues
AI workflows become operationally difficult when one user request expands into minutes or hours of work. A model may call tools, wait for external systems, request human approval, hit a rate limit, or resume after infrastructure failure. If all of that logic lives inside one short-lived request handler, a restart can lose progress and force expensive work to begin again. Azure now provides two complementary building blocks for this problem. Durable Task gives long-running AI and agent workflows checkpointing, state persistence, automatic recovery, and distributed coordination. Messaging services such as…
Microsoft AI-103: Designing AI Evaluation Datasets
An AI evaluation dataset is a product specification written as examples. It records the kinds of requests the system must handle, the evidence or outcomes that matter, and the failures that should block a release. If the dataset contains only easy prompts, evaluation will certify a system that works only when users behave exactly as expected. Microsoft Foundry supports reusable evaluation datasets in formats such as JSONL and CSV, versioned datasets for repeated evaluation runs, synthetic data generation, and evaluation of production traces. That makes the dataset more than a…