Practice Exams:

Microsoft AB-100: Microsoft Foundry Agent and Model ALM

AI & Machine Learning

An agent release is more than replacing one prompt with another. Business behavior can change when instructions, retrieval indexes, tools, model deployments, data schemas, credentials, and external policies drift independently. AB-100 architects must design application lifecycle management for agents, models, knowledge, and enterprise integrations. Microsoft Foundry provides components for code-first agent work, but the release obligation remains with the team: every promoted version needs a known configuration, representative tests, security checks, and a recovery path. This guide frames those dependencies as a reproducible release contract rather than a collection of environment screenshots.

On this page
  1. Treat an agent as a versioned dependency graph
  2. Separate development, validation, and production environments
  3. Version grounding data as carefully as prompts
  4. Establish evaluation gates before deployment
  5. Design rollback and incident response as part of ALM
  6. Put post-release monitoring into the ownership model

Treat an agent as a versioned dependency graph

Record the agent's instructions, tool contracts, knowledge sources, model deployment, safety configuration, and downstream APIs in a single release manifest. A model version alone is not enough: the same model can behave differently when a SharePoint knowledge source changes or a connector adds a new action. Assign owners and change controls to dependencies the agent team does not own directly.

For high-impact business workflows, capture the tool schema and expected permissions at deployment. A new connector field can introduce a failure even if prompt tests still pass. Keep example inputs and representative tool responses with the release package so future investigators can reproduce the behavior without relying on a vendor demo.

Separate development, validation, and production environments

The environment plan should keep test data and credentials separate from live customer records. Use controlled data samples that represent realistic format variability without exposing sensitive production documents unnecessarily. Promotion should carry only approved configuration and versioned artifacts. Secrets belong in an appropriate secret store and are injected at deployment, not copied into prompt files or release archives.

Define who can publish a new agent, who can update a model endpoint, who can approve a policy change, and who owns the final go/no-go decision. Segregation of duties matters when agents can take action in financial or regulated applications. A single developer's successful local run is not sufficient operational approval.

Version grounding data as carefully as prompts

Knowledge quality depends on document authority, effective dates, access labels, chunking, and retrieval indexes. Updating a corpus can change answers without changing code. Treat ingestion configuration, source manifests, access trimming, and index schema as release inputs. When a document is withdrawn, specify how quickly it must stop informing responses and how cached retrieval results are invalidated.

Create test cases for stale and revoked material. The model should abstain or refer the user to a current source when evidence is missing. The knowledge-sources guide discusses the permission and freshness questions that carry across platforms.

A test result loses its meaning if the document index changes without an accompanying release record. Capture the source collection, ingestion pipeline configuration, embedding model, chunking rules, filter policies, and freshness window used during evaluation. Sensitive data should not be copied into release artifacts; store identifiers and provenance instead. In production, a removed or reclassified document must stop influencing answers according to the organization’s deletion and permission-change expectations.

Establish evaluation gates before deployment

A credible release has deterministic contract tests, grounded answer evaluations, multi-turn conversation scenarios, adversarial inputs, and tests for each mutating tool. Include access-denied cases and simulations where a connector response is lost after a write. Track changes in quality, latency, cost, and error rates relative to a documented baseline. If a release changes an agent's allowed actions, require a separate authorization review rather than treating it as a minor wording update.

The evaluation data should include not only normal successes but examples of intentionally refused tasks. A model that becomes more helpful by bypassing restrictions is not improved. Keep the AB-100 testing guide in the regression plan even when the implementation is code-first.

Design rollback and incident response as part of ALM

Rollback can be difficult when a new agent version has already changed business records. Reverting prompts does not undo a payment, email, or case update. For consequential operations, combine reversible release configuration with idempotent write APIs, audit trails, compensating workflows, and human incident authority. Keep a documented kill switch for dangerous tools and a way to route requests to a safe read-only mode.

When an incident occurs, the team should identify the deployed agent configuration, source data versions, model choices, tool calls, policy decisions, and user identity. Logs must be appropriately secured themselves: diagnostic traces can contain sensitive text and should not become a second uncontrolled data store.

Rollback is not always a simple model-version switch. A new agent tool may have written records, a schema migration may have changed the knowledge store, and a safety rule may have been enforced in a separate service. Document which parts are reversible and which require compensating transactions. The runbook should allow operators to disable a high-risk tool or route traffic to a known-safe path before a complete redeployment. Preserve trace and evaluation evidence so the team can isolate whether a regression came from model behavior, retrieval, or integration.

Put post-release monitoring into the ownership model

Monitor groundedness, tool success, error taxonomy, denied operations, response latency, spend, and customer escalation. Avoid treating a fluent answer as proof of a successful downstream write. Reconcile tool results with system-of-record state and investigate unusual cost or action patterns that may signal prompt injection, bad routing, or an integration loop.

Every release should end with an owner, an alert recipient, and a scheduled review. The architectural goal is not to eliminate change but to make change observable, testable, and recoverable. That is the AB-100 ALM standard for a business AI service expected to survive outside the lab.

Related Posts

• Microsoft AI-300: Reproducibility Is the First Test of Production ML

• Microsoft AI-103: Safe Tool-Using Agents

• Microsoft AB-900: The Microsoft 365 Admin Role Now Includes AI

• Microsoft AI-103: Choosing Embeddings for Azure Search

• Microsoft AI-103: Diagnosing Hallucinations in Azure AI

• Microsoft AI-103: Reliable Python SDK Patterns for Azure AI

• Microsoft AI-103: Tool Calling for Azure AI Agents

• Microsoft AB-100: Building an AI Champions Program

• Microsoft AB-100: GitHub Copilot Context Engineering

• Microsoft AI-103: Text Analytics and Structured Extraction