AI & Machine Learning
Microsoft AI-103: Hybrid Search in Azure AI Search
Hybrid search is one of the most practical retrieval patterns in Azure AI Search because enterprise questions rarely fit neatly into either keyword search or vector search. Users mix exact identifiers with natural language. They search for product codes, policy names, dates, abbreviations, error messages, and concepts that may be phrased differently in the source material. A retrieval system that commits to only one matching strategy gives up useful evidence before ranking even begins. Azure AI Search runs full-text and vector queries in parallel inside a single hybrid request and…
Microsoft AI-103: Handling Hallucinations in Azure AI
Hallucination is a useful shorthand for a generated statement that is unsupported, fabricated, or inconsistent with the evidence the application should have used. It is not one failure with one fix. A bad answer can begin in retrieval, prompt construction, model behavior, tool output, stale data, or the application’s decision to answer when it should abstain. Azure provides controls at several layers. RAG can ground generation in enterprise evidence. Microsoft Foundry includes evaluators such as groundedness, relevance, response completeness, and task completion. Foundry tracing can connect model output to retrieval,…
Microsoft AI-103: Grounding Azure AI with Enterprise Data
Enterprise grounding is not simply connecting a model to documents. A production system has to find the right evidence, respect the caller’s permissions, preserve source attribution, handle changing data, and decide whether information should be indexed or queried live. The model only sees the result of that data architecture. Azure AI Search now underpins Foundry IQ, Microsoft’s managed knowledge layer for reusable, permission-aware knowledge bases in Microsoft Foundry. Azure AI Search also supports agentic retrieval with knowledge sources that can connect indexed search content and, in some cases, remote systems….
Microsoft AI-103: GenAIOps on Azure
GenAIOps is the operating discipline for generative AI systems after the prototype stage. It combines versioning, evaluation, tracing, monitoring, release control, cost management, and feedback so that model-driven behavior can change without becoming unpredictable. The focus is not one Azure product. It is the delivery loop that connects development evidence to production evidence. Microsoft Foundry now supports reusable evaluations, agent tracing with OpenTelemetry, Application Insights integration, monitoring dashboards, continuous evaluation, and analysis of deployed interactions. Azure DevOps or GitHub-based delivery pipelines can then use those signals to control promotion. The…
Microsoft AI-103: From AI Prototype to Production on Azure
An AI prototype proves that a capability is possible. Production proves that the capability can be trusted repeatedly under real users, real data, real load, changing models, operational failures, and governance requirements. The gap between those two states is where many AI projects become expensive: the demo worked, but the system was never designed to be operated. Microsoft Foundry now brings model deployment, agents, evaluation, tracing, monitoring, and governance closer together, while Azure services provide identity, networking, queues, search, storage, observability, and cost management. The production task is to turn…
Microsoft AI-103: Event-Driven AI Workflows on Azure
Event-driven architecture is a natural fit for AI work that begins when something changes rather than when a user waits on a synchronous request. A document arrives, a ticket changes state, a message is published, a file is written, or a business event crosses a threshold. Those events can trigger enrichment, classification, evaluation, retrieval updates, agent work, or downstream automation without holding the original caller open. Azure provides several services for different parts of that pattern. Event Grid routes events from sources to handlers. Service Bus queues and topics provide…
Microsoft AI-103: Durable AI Workflows with Queues
AI workflows become operationally difficult when one user request expands into minutes or hours of work. A model may call tools, wait for external systems, request human approval, hit a rate limit, or resume after infrastructure failure. If all of that logic lives inside one short-lived request handler, a restart can lose progress and force expensive work to begin again. Azure now provides two complementary building blocks for this problem. Durable Task gives long-running AI and agent workflows checkpointing, state persistence, automatic recovery, and distributed coordination. Messaging services such as…
Microsoft AI-103: Designing AI Evaluation Datasets
An AI evaluation dataset is a product specification written as examples. It records the kinds of requests the system must handle, the evidence or outcomes that matter, and the failures that should block a release. If the dataset contains only easy prompts, evaluation will certify a system that works only when users behave exactly as expected. Microsoft Foundry supports reusable evaluation datasets in formats such as JSONL and CSV, versioned datasets for repeated evaluation runs, synthetic data generation, and evaluation of production traces. That makes the dataset more than a…
Microsoft AI-103: Deploying Fine-Tuned Models on Azure
A fine-tuned model is not production-ready when training finishes. It becomes operational only after the team has selected a checkpoint, passed quality and safety evaluation, chosen a supported deployment type, configured access, tested inference behavior, and defined how the deployment will be upgraded or retired. Deployment is the point where a training artifact becomes a service with cost, capacity, and lifecycle responsibilities. Microsoft Foundry currently allows fine-tuned Azure OpenAI models to be deployed for inference after training. Deployment requires appropriate control-plane permissions, and supported deployment types vary by model and…
Microsoft AI-103: Cost Control for Azure AI Apps
Azure AI cost is not one number. A production application can generate charges from model tokens, provisioned throughput, fine-tuned model hosting, search, storage, Application Insights, API Management, functions, databases, and external tools. The only useful cost model is therefore a workload model: what a successful user task invokes, how often it happens, and how that behavior changes under load. Microsoft Foundry currently supports pay-as-you-go model usage, provisioned throughput, fine-tuned model hosting, and gateway-level token controls. Azure Cost Management provides the billing view, while model and application telemetry explain why usage…
Microsoft AI-103: Chunking Strategies for Azure RAG
Chunking determines what a retrieval system is capable of finding. In an Azure RAG application, documents are rarely useful as one large block of text. The retrieval layer needs units that are small enough to match a focused question but large enough to preserve the context that makes the answer meaningful. That balance is why chunking deserves to be designed and evaluated rather than treated as an indexing default. Current Azure AI Search guidance supports several approaches, from fixed text splitting to structure-aware chunking with document layout information. Integrated vectorization…
Microsoft AI-103: Choosing Embeddings on Azure
Embedding choice affects retrieval quality, index size, query latency, re-indexing cost, and the long-term shape of a RAG system. It is easy to treat embeddings as an implementation detail because they sit behind vector search, but the vector dimensions are part of the index schema and the model determines how text is represented. Changing either later can require rebuilding the vector corpus. Azure OpenAI currently supports third-generation embedding models such as text-embedding-3-small and text-embedding-3-large, alongside older options in some services. Azure AI Search can use these models through integrated vectorization…
Microsoft AI-103: Choosing Azure AI Deployment Models
Azure AI deployment choices determine more than where a model runs. They affect where inference data can be processed, how capacity is allocated, whether billing is pay-per-token or reserved, how much latency variation to expect, and whether the workload is suited to real-time or asynchronous processing. Choosing the wrong deployment type can create a compliance or reliability problem even when the model itself is a good fit. Microsoft Foundry currently uses serverless API as the preferred deployment option for a broad set of Foundry Models. Within that option, deployments can…
Microsoft AI-103: Capacity Planning for Azure AI
Capacity planning for Azure AI begins with a simple correction: serverless does not mean unlimited. Model APIs can have tokens-per-minute limits, requests-per-minute limits, concurrency constraints, regional availability, deployment-specific quotas, and capacity that changes by model and SKU. A workload that looks small in request counts can still be large in tokens, while a high-request workload with tiny prompts can hit request limits before token limits. Microsoft Foundry separates standard pay-per-token deployment types from provisioned throughput and batch options. Azure OpenAI quotas are also scoped by factors such as subscription, region,…
Microsoft AI-103: Canary Releases for AI Models
A canary release exposes a new AI version to a deliberately small portion of production traffic before broader rollout. The idea is familiar from ordinary software delivery, but AI adds a second dimension: the endpoint can remain technically healthy while answer quality, safety behavior, tool selection, or retrieval performance gets worse. The canary therefore needs behavioral release criteria as well as infrastructure metrics. Azure Machine Learning managed online endpoints support percentage-based traffic routing between deployments, which makes them a natural fit for canary rollout. For model services where native traffic…