Practice Exams:

Microsoft AI-103: Handling Hallucinations in Azure AI

Hallucination is a useful shorthand for a generated statement that is unsupported, fabricated, or inconsistent with the evidence the application should have used. It is not one failure with one fix. A bad answer can begin in retrieval, prompt construction, model behavior, tool output, stale data, or the application’s decision to answer when it should abstain.

Azure provides controls at several layers. RAG can ground generation in enterprise evidence. Microsoft Foundry includes evaluators such as groundedness, relevance, response completeness, and task completion. Foundry tracing can connect model output to retrieval, tools, tokens, and latency. Azure AI Content Safety also offers groundedness-related capabilities in supported scenarios.

The practical goal in Azure AI engineering is not to promise that hallucinations disappear. It is to reduce the conditions that produce them, detect unsupported output, and make failures diagnosable.

Find where the unsupported claim entered the pipeline

A final answer can be wrong because the correct evidence was never retrieved. It can be wrong because the evidence was retrieved but ranked below the context cutoff. It can be wrong because the model ignored or misread good evidence. It can also be wrong because a tool returned stale or malformed data.

Those failures require different fixes. Changing the prompt will not repair a missing document. Replacing the embedding model will not fix an incorrect API response.

Hallucination tracing is therefore the right diagnostic mindset: locate the stage where support was lost before tuning anything.

Ground the model when the answer depends on changing facts

Foundation models are not databases for private or frequently changing enterprise information. Use retrieval or tools when the answer depends on current internal facts. Azure AI Search can provide indexed RAG and agentic retrieval, while Foundry IQ can organize permission-aware knowledge sources for agents.

The system should pass only evidence relevant to the question and authorized for the caller. More context is not automatically safer; irrelevant or conflicting text can confuse generation.

Enterprise grounding is the architecture layer that makes factual enterprise answers possible.

Evaluate retrieval before evaluating the prose

If the correct passage is not in the retrieved set, the generator is being asked to answer without the evidence it needs. Build a retrieval benchmark with known relevant passages and measure whether they appear high enough in the result set to enter the prompt.

Chunking, embeddings, hybrid search, metadata filters, and semantic ranking all affect this stage. RAG chunking and hybrid search should be evaluated independently from the generation model.

This separation prevents teams from using a larger model to hide a retriever that is systematically missing the source of truth.

Groundedness measures support, not universal truth

Foundry’s groundedness evaluator measures whether a response is supported by the provided context. That is valuable for RAG because the application knows which evidence it supplied. It does not prove that the source itself is correct or current.

A response can be perfectly grounded in a stale policy document. That makes source governance and freshness part of hallucination control. The system needs both trustworthy evidence and generation that stays within that evidence.

Use groundedness as one metric beside relevance, completeness, task success, and domain-specific checks rather than as a universal truth score.

Teach the system when to abstain

An AI application should not be forced to answer every question. If retrieval returns no strong evidence, the safer behavior may be to say the information is unavailable, ask for clarification, or direct the user to an authoritative source.

Abstention should be tested. Include evaluation cases where the correct behavior is not to answer. Measure whether the system invents a response under pressure.

This is particularly important for high-risk workflows where a confident unsupported answer can lead to a real action rather than merely a bad conversation.

Use tool outputs as evidence with validation

Agents can reduce hallucination by calling tools for current facts, but tools create their own failure modes. An API can return an error page, partial result, stale cache, or value in an unexpected unit. The model may then confidently interpret bad data.

Validate tool responses before passing them to the model. Use typed schemas, status checks, range checks, and explicit error states. Tool-call success is different from business correctness.

The same logic appears in agent testing: the model should be evaluated on how it chooses and uses tools, not only on the final prose.

Trace production behavior so failures become evidence

Foundry tracing can record the request path through model calls, retrieval, tools, timing, and token use. Application Insights then provides a place to investigate production interactions. This is essential when a hallucination appears only under one document, one model version, or one tool response.

AI observability should preserve correlation across those stages while minimizing sensitive content. Operators need enough evidence to reconstruct why the answer was produced.

When a failure is understood, create a sanitized representative case and add it to the evaluation dataset so the regression can be caught before the next release.

Release model and prompt changes through regression tests

A new model can reduce one type of hallucination and increase another. A prompt intended to make answers more helpful can encourage the system to guess. A retrieval change can improve recall while introducing conflicting sources.

Run the stable dataset from evaluation datasets against candidate changes. Compare groundedness, task success, abstention, latency, and cost. Then use controlled rollout rather than assuming offline gains will transfer perfectly to production.

This is where GenAIOps closes the loop: production failures update the test suite, and the test suite controls future releases.

Hallucination control is a system design problem

No single prompt, evaluator, or safety filter can guarantee factual output. Reliable systems combine authoritative data, good retrieval, validated tools, explicit abstention, evaluation, tracing, and release discipline. The model is one participant in that system.

For the current Azure AI certification role, the practical skill is to diagnose unsupported answers by layer. Fix the source, retriever, tool, prompt, model, or policy that actually failed. That is more effective than treating every wrong answer as the same mysterious model behavior.

Unsupported answers should also be categorized rather than counted as one generic defect. A system can fabricate an entity, use a real source incorrectly, state an outdated fact, overgeneralize beyond the evidence, or invent a tool result. Those categories point to different owners and fixes. A retrieval team can act on missing evidence; an application team can fix stale tool data; a prompt or model change may address unsupported extrapolation.

Keep a small incident taxonomy in production quality review. When a sampled answer fails, assign the failure to the earliest layer that made the final answer unreliable. Over time, that distribution shows whether engineering effort should focus on ingestion, ranking, tool validation, prompt design, or model choice.

Do not hide uncertainty merely to make the assistant sound confident. When evidence is partial, the response can state the boundary of what is supported and identify what would be needed to answer more strongly. In enterprise settings, a calibrated answer with clear limits is often more useful than a fluent answer that fills every gap.

Finally, evaluate hallucination controls under adversarial and degraded conditions. Remove the expected source, inject conflicting documents, return a tool timeout, and lower retrieval quality deliberately. The system should fail in a predictable way. Reliability is demonstrated not only when the correct evidence is present, but also when the application behaves responsibly when that evidence is missing.

Metrics should distinguish frequency from severity. Ten harmless wording mistakes are not equivalent to one fabricated instruction that triggers an operational action. Weight quality review by consequence and user exposure so teams do not optimize a headline hallucination rate while missing a rare but dangerous failure.

Escalation paths matter for recurring uncertainty. If the application repeatedly abstains on the same high-value topic, that may indicate missing source coverage rather than a model problem. Feed those patterns back to content owners and retrieval engineering so the system’s knowledge boundary improves over time instead of merely producing better refusals.

When the same unsupported pattern survives several model or prompt changes, revisit the architecture. Persistent hallucination can indicate that the task itself requires retrieval, a deterministic calculation, or a validated business rule rather than more generative reasoning. The safest fix is sometimes to remove responsibility from the model.

Track those architectural fixes separately from prompt tuning so the team can see which interventions actually reduce recurrence.

Related Posts

• Microsoft AI-103: Azure AI Search for RAG

• Microsoft AI-103: Blue-Green Releases for AI Endpoints

• Microsoft AI-103: Building Multi-Agent Workflows on Azure

• Microsoft AI-103: Canary Releases for AI Models

• Microsoft AI-103: Capacity Planning for Azure AI

• Microsoft AI-103: Durable AI Workflows with Queues

• Microsoft AI-103: Event-Driven AI Workflows on Azure

• Microsoft AI-103: From AI Prototype to Production on Azure

• Microsoft AI-103: GenAIOps on Azure

• Microsoft AI-103: Grounding Azure AI with Enterprise Data