Microsoft AI-103: Online Evaluation for AI Systems
Online evaluation closes the gap between a controlled benchmark and the behavior users actually experience. Prelaunch datasets are essential for release gates, but production traffic introduces new phrasing, unexpected tool paths, changing knowledge, rare edge cases, and interactions the development team never wrote down. A mature AI system therefore needs a way to measure deployed behavior without treating every quality issue as a support anecdote.
Microsoft Foundry can evaluate deployed interactions and Application Insights traces with the same evaluation framework used for offline quality analysis. Some production-trace and conversation evaluation capabilities are currently preview, so teams should distinguish experimentation from service-level commitments. The useful architectural idea is stable even while individual features mature: production traces should become measurable evidence.
That makes online evaluation part of Azure AI engineering, not a separate analytics project.
Offline benchmarks and online evaluation answer different questions
Offline evaluation asks whether a known candidate behaves well on a controlled set of cases. Online evaluation asks what the deployed system is actually doing across real traffic.
The first supports release decisions. The second supports monitoring, incident investigation, and discovery of new failure modes. One cannot replace the other.
Evaluation datasets give the team a stable regression set, while production traces reveal whether that set still represents the workload.
Trace evaluation avoids replaying production requests
Foundry can evaluate interactions already captured in Application Insights by using trace data. This is valuable because the system can score what actually happened rather than rerunning a request against a potentially different model, prompt, knowledge base, or external tool state.
Trace evaluation can target specific operation IDs or discover recent traces for an agent, depending on the scenario. That makes it useful for investigating a known incident or sampling live behavior.
The production trace should therefore preserve enough context to reconstruct the relevant request path without retaining unnecessary sensitive content.
Conversation-level evaluation matters for agents
A single-turn answer can look correct while a multi-turn agent fails the overall task. The agent may forget a constraint, repeat a question, misuse a tool, or lose state across handoffs.
Conversation-level evaluation considers the interaction as a sequence. Current Foundry capabilities for deployed conversations are preview, but they point toward a stronger production model: evaluate whether the whole task succeeded, not only whether one response sounded coherent.
Agent testing should use the same principle before and after deployment.
Sampling is necessary for cost and signal quality
Evaluating every production interaction with several model-assisted metrics can be expensive. It can also generate large volumes of low-value data when most traffic is routine.
Use sampling that reflects product risk. High-impact tool actions may justify dense evaluation. Routine low-risk chat may be sampled. Known incident cohorts can receive targeted evaluation until confidence is restored.
The sampling strategy should preserve important languages, task types, customer segments, and rare but consequential workflows rather than selecting only the easiest high-volume traffic.
Production evaluation needs clear metrics
Choose metrics based on the system. RAG may need groundedness, relevance, citation quality, and completeness. Agents may need task completion, tool-call accuracy, unnecessary-action rate, or policy compliance. Structured extraction may need schema validity and field accuracy.
Do not reduce every product to one aggregate score. Critical dimensions should have their own alerting or review thresholds.
Hallucination control is stronger when unsupported answers are measured separately from other quality defects.
Use production failures to update the benchmark
A real failure is valuable because it reveals a gap in the prelaunch evaluation set. Once the issue is understood, create a safe representative case and add it to the regression suite.
Foundry can convert selected traces into a curated, versioned dataset. That closes the loop between observability and predeployment testing.
GenAIOps should make this routine: observe, diagnose, convert the failure into evidence, fix the system, and make the next release prove that the failure stays fixed.
Online evaluation should be version-aware
Production traffic often spans more than one model, agent version, prompt version, or rollout cohort. Evaluation results are useful only when they can be attributed to the behavior package that produced them.
Tag traces with deployment, model, agent version, prompt version, retrieval index, and application build where practical. If a canary is live, compare baseline and candidate cohorts instead of pooling their scores.
This is where online evaluation becomes release evidence rather than general quality telemetry.
Quality alerts should lead to traces, not guesswork
An alert that says groundedness fell is a starting point. Operators need to inspect representative traces, retrieved evidence, tool outputs, latency, and version changes to determine the root cause.
AI observability should connect the quality signal to the request path. A retriever regression and a model regression can produce the same bad answer but require different fixes.
Quality monitoring is therefore strongest when evaluation and tracing live in the same operational workflow.
Keep preview status explicit in production planning
Some current Foundry capabilities for trace and deployed-conversation evaluation are preview. Preview features can be excellent for testing an operating model, but they should not be represented internally as having the same support or SLA expectations as generally available features.
For the current AI-103 ecosystem, the durable skill is to combine stable offline evaluation with measured production behavior. The tool surface will evolve, but AI teams will continue to need the same loop: benchmark before release, observe after release, turn real failures into reusable tests, and use evidence to control the next change.
Privacy and retention have to be designed into the evaluation path. Production traces can contain user prompts, retrieved text, tool arguments, and outputs. The evaluation system should minimize or transform sensitive content where possible and should not create a second uncontrolled archive simply because quality teams want examples. Sampling and redaction rules belong in the operational design.
Human review remains useful for metrics that are subjective or business-specific. Model-assisted evaluators can scale, but reviewers should periodically inspect disagreements and calibrate the rubric. If a human reviewer and automated evaluator consistently diverge on one scenario class, the metric may need better instructions or a different evaluator.
Online evaluation should also distinguish discovery from alerting. Broad sampled evaluation can discover emerging failure patterns, while a narrower high-confidence metric may be appropriate for automated release rollback or incident alerts. Not every quality score is stable enough to drive an automatic operational action.
Keep trend views version-aware. If a new agent version receives ten percent of traffic, an aggregate quality chart can hide whether the decline belongs to the candidate or to the baseline. Cohort labels make it possible to compare like with like during gradual rollout.
The strongest online evaluation program ends by changing the offline suite. A production issue that matters should become a durable test, while obsolete or redundant cases should be reviewed periodically. The benchmark should evolve with the product rather than grow indefinitely without curation.
Evaluation cadence should reflect how quickly the system changes. A stable internal assistant may need periodic sampled review, while an actively changing agent under canary release may justify continuous evaluation of the candidate cohort. The cadence should increase around releases and incidents, then return to a sustainable baseline.
Cost estimates from evaluator runs should be tracked separately from product inference cost. Model-assisted evaluation consumes tokens and compute even when it does not affect the user request. This is part of the operating budget and can become significant at scale.
Quality review should also preserve the raw evidence needed for appeal. If an automated evaluator marks an interaction as poor, a reviewer should be able to inspect the response, relevant context, and trace metadata before deciding whether the score indicates a real defect.
Evaluation jobs should have owners and review queues. A sampled interaction that fails a metric needs a path to triage: true product defect, evaluator error, unsupported use case, data issue, or acceptable edge behavior. Without that classification, evaluation accumulates scores without producing engineering decisions.
Use baseline comparisons around releases. Foundry can compare evaluation runs, and production cohorts can be evaluated against the same rubric. The important habit is to define the reference version before rollout so teams do not search for a favorable baseline after seeing the candidate results.
Document evaluator limitations in the runbook. Some metrics require reference context, some depend on model-assisted judgment, and some are more reliable for one task family than another. Operators should know which scores are suitable for trend monitoring, which can block a release, and which always need human review.