Practice Exams:

Microsoft AI-103: Prompt Versioning in Azure AI

Prompt versioning is necessary because prompt text is production behavior. A change to system instructions, examples, tool guidance, or retrieval framing can alter quality, safety, latency, cost, and tool use without any application-code change. Teams that edit prompts directly in a portal without preserving history lose the ability to explain why behavior changed or to restore the last known-good configuration.

Current Microsoft Foundry provides immutable versions for prompt agents: after a saved version is created, later edits are saved as a new version. Foundry’s code-first path also allows agent definitions to live in SDK or REST workflows that can participate in source control and CI/CD. Older Prompt flow variants remain useful context for legacy projects, but Microsoft has announced Prompt flow retirement for April 20, 2027 and does not recommend it for new development.

That makes versioning a current Azure AI engineering responsibility rather than a documentation preference.

Version the whole instruction set

A prompt is more than one text box. Behavior can depend on system instructions, examples, variables, model parameters, tool descriptions, response schemas, and retrieval instructions.

Keep those elements together as one versioned configuration. A change to temperature or tool instructions can be just as meaningful as a wording change.

Prompt layers should be visible in the version record so reviewers know which part of the instruction hierarchy changed.

Use immutable versions for released agents

Foundry prompt-agent versions are immutable after save. That is a strong production property because a conversation, trace, or evaluation can point to a specific agent version that will not silently change later.

Use unsaved drafts for experimentation, then save a new version when the candidate deserves persistent evaluation or release consideration.

Do not overwrite the production meaning of an old version merely to keep the version list short.

Keep source control as the durable history

Portal versioning is useful, but code-first definitions make change review stronger. Store instructions and configuration in source control when the project supports that workflow. Use pull requests to review meaningful behavior changes.

This creates a durable history outside one portal experience and lets infrastructure, application code, and agent configuration move through coordinated CI/CD.

Prompt management becomes manageable when prompts are treated like code assets rather than copy pasted text.

Version model choice beside the prompt

The same prompt can behave differently when the model version changes. Foundry model deployments can have explicit model versions and upgrade policies, so prompt evaluation should record the exact deployment version.

A prompt version without the model context is incomplete release evidence.

For high-control workloads, pin model versions until the new version has been tested against the evaluation suite.

Attach evaluation results to the candidate

A prompt should not be promoted because it “looks better” in the playground. Run it against a stable evaluation set and compare it with the current baseline.

Evaluation datasets allow the same scenarios to be reused across prompt and agent versions. Store the result or run identifier with the release record.

If the candidate improves one metric but regresses a critical slice, keep the existing version.

Use promotion instead of direct replacement

Prompt optimization and testing should create a candidate, then promote that candidate to a new version only after review. Current Foundry prompt-agent optimization follows that pattern: candidates can be compared with a baseline and promoted as a new agent version.

Promotion is a decision point. It should include quality, safety, latency, cost, and operational evidence appropriate to the workload.

This creates a clear separation between experimentation and production state.

Rollback should restore the complete behavior package

Rolling back only the text prompt may not restore the previous behavior if the model, tools, retrieval settings, or application code changed too.

GenAIOps should define the behavior package that moves together through release and rollback.

Keep the previous known-good version available until the new candidate has completed its soak period.

Treat Prompt flow variants as legacy migration input

Classic Prompt flow supported variants for comparing prompt text and connection settings. That concept remains useful, but the current product direction matters: Microsoft has announced Prompt flow retirement in Foundry and Azure Machine Learning for April 20, 2027.

Existing Prompt flow workloads should plan migration rather than build new long-term dependencies on the classic experience.

New prompt-agent work should use current Foundry agent versioning and supported development paths.

Make versioning visible in production traces

Telemetry should identify the prompt or agent version that produced an interaction. Without that field, a quality regression can be difficult to correlate with a change.

Online evaluation becomes far more useful when traces can be grouped by version and compared across baseline and candidate cohorts.

For current Azure AI teams, prompt versioning is therefore a release-control discipline: immutable versions, source history, model context, evaluation evidence, controlled promotion, and traceable rollback.

Naming conventions can make version history easier to use. A version number should be accompanied by release metadata such as purpose, change summary, model deployment, evaluation run, and owner. An immutable version is valuable only if future operators can tell why it existed and whether it was ever promoted.

Separate experimentation from release candidates. Developers may try many local or draft prompt changes. Only candidates that pass basic tests need durable production-facing versions. This keeps the release history meaningful without preventing rapid iteration.

Prompt variables also deserve contracts. A template can be versioned correctly while a calling application sends a different field, empty value, or untrusted string into a high-impact placeholder. Validate variable types and required fields so prompt versioning does not create false confidence about runtime behavior.

Model upgrade policies should be reviewed with prompt versions because an automatically upgraded model can change behavior underneath an unchanged prompt. Where behavior stability matters, pin or control model upgrades and evaluate the new model version before changing the deployment.

Finally, define how old prompt versions are retired. Keep versions needed for audit and rollback, but remove obsolete drafts from active deployment paths. A clean version history reduces the chance that a stale configuration is accidentally reactivated during an incident.

Version identifiers should appear in logs, traces, and support tools so a user-reported issue can be tied to the exact configuration. If support cannot tell which prompt version handled the conversation, the version history provides little operational value.

Evaluation data should also be versioned independently. A prompt may improve because the test set changed rather than because the instructions improved. Store the dataset version and evaluator configuration with each comparison so the result can be reproduced.

Prompt versioning becomes especially important when several teams share an agent. Establish one promotion owner or a clear approval workflow so two parallel edits do not both become “the next version” without comparison. Branching experimentation is healthy; ambiguous production ownership is not.

For locally stored prompts or code-defined agents, use the same release discipline even if the platform does not create a portal version automatically. Git history, tags, evaluation artifacts, and deployment metadata can provide the immutable reference needed for production.

Release notes should focus on behavioral intent. Record whether the version changes tone, refusal policy, tool selection, retrieval instructions, output schema, or task scope. That is more useful during regression analysis than a generic note such as “prompt cleanup.”

Versioning also helps experimentation stay honest. When a team compares two prompt variants, save the exact text and parameters that produced each result. Copying the “winner” into another editor by hand can introduce small changes that make the recorded evaluation no longer representative of the deployed version.

For shared libraries of prompts, establish dependency tracking. If several agents import the same instruction component, changing that component can affect multiple products. Version the shared element and require consumers to update deliberately rather than receiving an invisible global prompt change.

Prompt ownership should be explicit as well. Product, safety, and engineering teams may all propose instruction changes, but one accountable owner should decide what becomes production behavior. Clear ownership keeps urgent edits from bypassing evaluation and prevents conflicting prompt changes from being merged without understanding their combined effect.

Version tags can also simplify support: a conversation report that includes agent version, model version, and application build gives responders an immediate starting point.

Periodic cleanup should archive superseded drafts without erasing the production versions needed for rollback, audit, and historical comparison.

Keep the release record concise enough that engineers will actually maintain it; a small reliable history is better than a complex version process people bypass.

Related Posts

• Microsoft AI-103: Agent Identity in Azure AI Foundry

• Microsoft AI-103: Capacity Planning for Azure AI

• Microsoft AI-103: Choosing Azure AI Deployment Models

• Microsoft AI-103: Choosing Embeddings on Azure

• Microsoft AI-103: Chunking Strategies for Azure RAG

• Microsoft AI-103: GenAIOps on Azure

• Microsoft AI-103: Grounding Azure AI with Enterprise Data

• Microsoft AI-103: Handling Hallucinations in Azure AI

• Microsoft AI-103: Hybrid Search in Azure AI Search

• Microsoft AI-103: Latency Tuning for Azure AI Apps