Practice Exams:

Microsoft AI-103: Designing AI Evaluation Datasets

An AI evaluation dataset is a product specification written as examples. It records the kinds of requests the system must handle, the evidence or outcomes that matter, and the failures that should block a release. If the dataset contains only easy prompts, evaluation will certify a system that works only when users behave exactly as expected.

Microsoft Foundry supports reusable evaluation datasets in formats such as JSONL and CSV, versioned datasets for repeated evaluation runs, synthetic data generation, and evaluation of production traces. That makes the dataset more than a one-time test file. It can become the stable regression suite used across models, prompts, agents, and releases.

For Azure AI engineering, the challenge is to build a dataset that represents the actual workload without becoming a frozen snapshot that ignores new production behavior.

Begin with decisions, not evaluator names

Before choosing groundedness, relevance, task completion, or another metric, decide what product decision the evaluation must support. A prelaunch test may ask whether a system is safe enough to expose to users. A model comparison may ask which candidate achieves the best quality within a latency budget. A regression suite may ask whether a prompt change breaks existing tasks.

Each decision needs different examples. A dataset for grounded RAG needs queries with source context and expected evidence. A tool-using agent needs scenarios where the correct action is known. A conversational assistant needs multi-turn cases where state and follow-up matter.

The article on agent testing shows why evaluating only the final answer can miss failures in the behavior that produced it.

Represent the real distribution and the risky tail

A useful dataset has common cases because common cases dominate user experience. It also has rare but consequential cases because those failures dominate operational risk. Do not sample production traffic blindly and assume frequency alone defines importance.

Stratify the dataset by task type, user intent, language, document source, tool path, risk category, and other dimensions that matter to the product. Add difficult examples near policy boundaries, ambiguous requests, incomplete information, and known historical failures.

This gives the evaluation enough structure to detect when an overall score improves while one critical cohort becomes worse.

Keep training data and evaluation data meaningfully separate

A model or prompt should not be optimized directly against every evaluation example until the suite becomes another training set. Hold back cases that the development loop does not see. For fine-tuning, use an evaluation set that was not included in training. For prompt work, keep part of the regression suite protected from repeated manual tuning.

Complete isolation is difficult in modern AI systems, especially with foundation models trained on broad public data. The practical goal is process separation: do not intentionally optimize on the exact examples used to make the final release decision.

fine-tuning data helps keep that boundary visible because the team can prove which examples shaped training and which remained evaluation-only.

Ground truth should match the task

Not every evaluation needs one perfect reference answer. Extraction tasks may have exact expected fields. Classification can have a correct label. RAG may need a set of relevant source passages rather than one wording of the answer. Open-ended assistance may need a rubric describing what a good response must include.

Use exact ground truth where the task is objective. Use structured criteria where multiple valid outputs are possible. Do not force generative output into brittle string matching when semantic quality is what matters.

Foundry evaluators can map dataset fields to different quality checks, which is useful only when the dataset schema contains the information those evaluators need.

Synthetic data is a bootstrap, not a substitute for reality

Foundry can generate synthetic evaluation queries, which is useful before launch or when production traffic is sparse. Synthetic examples can expand coverage across scenarios the team already understands. They can also repeat the assumptions and blind spots of the model used to generate them.

Review synthetic examples before treating them as a benchmark. Mix them with curated domain cases and later with sanitized production traces. Track the origin of every example so analysts can see whether a score is being driven by synthetic or real behavior.

The warning in synthetic data applies directly to evaluation datasets.

Production traces should feed the regression suite

Foundry can evaluate deployed interactions and can convert trace-derived behavior into reusable evaluation data. That creates a valuable feedback loop: production reveals a new failure, the team investigates it, a safe representative case is added to the dataset, and future releases are tested against it.

Do not copy sensitive prompts or customer data into a long-lived benchmark without governance. Redact, synthesize, or transform production examples while preserving the failure pattern that matters.

This workflow connects evaluation to AI observability. Traces explain what happened; the evaluation dataset prevents the same class of failure from returning unnoticed.

Version the dataset and preserve stable slices

Evaluation datasets should evolve, but scores are only comparable when the underlying test set is known. Give datasets versions. Keep a stable core slice for longitudinal comparison and add new slices for emerging product behavior.

When the dataset changes significantly, record both old and new scores instead of pretending the metric is directly continuous. A harder test set can make a better system appear to regress.

Versioned Foundry datasets are useful for CI/CD because the pipeline can declare exactly which benchmark version a candidate must pass.

Use multiple metrics instead of one headline score

Quality is multidimensional. Foundry includes evaluators for task completion, coherence, groundedness, relevance, response completeness, safety, and agent-tool behavior. A single averaged score can hide a severe failure in one dimension.

Define release gates by the product’s priorities. A RAG assistant may require groundedness above a threshold and zero regression on a sensitive domain slice. An action-taking agent may require tool-selection and tool-input accuracy before conversational style matters.

These gates should be understandable to engineers and product owners. If nobody can explain why the system passed, the evaluation pipeline is producing ceremony rather than evidence.

Make the dataset part of release engineering

A good evaluation dataset survives individual model choices. It can compare a new model, prompt, retriever, agent graph, or fine-tuned checkpoint using the same business expectations. That makes it one of the most reusable assets in the AI delivery pipeline.

For the current Azure AI certification role, evaluation is not an optional quality pass after implementation. It is how architecture changes become measurable. A dataset should tell the team what must remain true as the AI system changes.

Give every case a stable scenario identifier and enough metadata for slicing. Useful fields can include task family, risk level, source domain, language, expected tool, expected source, difficulty, and whether the example came from curation, synthetic generation, or production. The exact schema varies, but the purpose is the same: when a score moves, the team should be able to identify which kinds of behavior changed.

Dataset balance also needs deliberate review. A benchmark with hundreds of routine prompts and two high-risk action scenarios can produce an excellent average while still missing the behavior the organization cares about most. Critical slices should have their own gates when necessary. A release might need to pass an overall quality threshold and separately pass a security-sensitive or high-value workflow slice.

An evaluation dataset should also document its limits. A finite suite does not certify all future behavior, model-assisted evaluators can make mistakes, and synthetic generation can repeat the assumptions of the generator. The dataset is strongest when it is treated as a maintained engineering asset with known coverage rather than as a one-time certificate that the AI system is correct.

Reviewer instructions should be versioned too. Human labels are only useful when reviewers apply the same rubric. Define what counts as a pass, what evidence can support a judgment, how ambiguous cases are handled, and when a reviewer can mark an example as invalid rather than forcing a score. Periodic agreement checks can reveal when the rubric itself has become unclear.

For automated evaluators, store the evaluator configuration and model version with the result. The same discipline applies to agent evaluation, where tool use and task behavior need repeatable criteria. A score can change because the system under test changed or because the judge changed. Without that provenance, longitudinal comparisons can become misleading.

Sampling strategy matters when production data is converted into evaluation cases. Avoid taking only the easiest high-volume interactions or only the incidents that reached support. A representative sample should include ordinary traffic, difficult edge cases, and important low-frequency workflows so the suite reflects both experience and risk.

Related Posts

• Microsoft AI-103: Agent Identity in Azure AI Foundry

• Microsoft AI-103: Azure AI Content Safety in Practice

• Microsoft AI-103: Building Multi-Agent Workflows on Azure

• Microsoft AI-103: Canary Releases for AI Models

• Microsoft AI-103: Capacity Planning for Azure AI

• Microsoft AI-103: Choosing Azure AI Deployment Models

• Microsoft AI-103: Choosing Embeddings on Azure

• Microsoft AI-103: Chunking Strategies for Azure RAG

• Microsoft AI-103: Cost Control for Azure AI Apps

• Microsoft AI-103: Deploying Fine-Tuned Models on Azure