Microsoft AI-103: Azure AI Foundry Model Selection
Model selection in Azure should begin with the workload, not the catalog. A model can score well on a public benchmark and still be a poor production fit because it is unavailable in the required deployment type, too slow for the interaction pattern, too expensive at expected token volumes, weak on the organization’s own data, or incompatible with a required tool or modality.
Microsoft Foundry gives teams access to models from Microsoft, Azure OpenAI, and multiple partner and community providers. That breadth is valuable, but it makes a disciplined selection process more important. The engineering goal is not to find the universally best model. It is to identify the smallest set of models that meets measurable quality, safety, operational, and commercial requirements for a specific task.
The current AI-103 study scope explicitly includes choosing appropriate models, deployment options, retrieval services, and evaluation approaches. That mirrors production reality: model choice is connected to architecture decisions across the whole Azure AI engineering stack.
Start with the job the model must perform
Define the task before comparing models. Classification, extraction, summarization, conversational assistance, code generation, multimodal analysis, tool calling, and long-form reasoning have different requirements. Even two chat experiences can require very different models if one needs strict structured output and the other needs open-ended analysis.
Write a short workload contract: expected inputs, output format, maximum acceptable latency, languages, context size, tool requirements, grounding strategy, safety constraints, and approximate request volume. Add a small set of failure conditions that are unacceptable. A model that violates a hard constraint can be removed before expensive evaluation begins.
This avoids a common mistake in foundation model selection: teams compare features in the abstract instead of deciding what the product actually needs.
Separate hard constraints from quality preferences
Some requirements are binary. A model must be available in a deployment type that satisfies data-processing policy. It must support the needed modality. It must work with the chosen API surface. It must be deployable in a region or data zone the organization can use. If a workload needs a specific structured-output or tool-calling capability, lack of that capability is not something a slightly better benchmark score can compensate for.
Other requirements are preferences that can be measured and traded off: answer quality, latency, cost, verbosity, robustness to noisy input, and consistency across repeated calls. Keeping the two groups separate reduces noise. First eliminate impossible candidates; then evaluate the viable ones.
Availability can also change over a model’s lifecycle. Foundry documentation distinguishes model versions, deployment types, and retirement timelines. Production designs should therefore record not just a model family but the deployed version and an upgrade plan.
Build an evaluation set from real work
Public benchmarks are useful for orientation, but the best model for your application is determined by your own prompts, documents, languages, output formats, and risk tolerance. Build an evaluation set from representative tasks. Include easy cases, difficult cases, long-context inputs, ambiguous requests, edge conditions, and examples that previously caused failures.
Use task-specific metrics. Extraction can be scored against known fields. Classification can use precision and recall. RAG can evaluate retrieval and grounded answer quality separately. Agent workflows can measure tool choice, task completion, unnecessary actions, and policy adherence. Human review still matters where the output is subjective, but reviewers need a rubric rather than a vague preference.
The evaluation set becomes a durable engineering asset. It can compare model candidates now and later detect regressions when a model version, prompt, retriever, or safety policy changes.
Latency and throughput change the practical ranking
Interactive systems care about more than average response time. Time to first token, full completion time, tail latency, request concurrency, and token throughput can all affect user experience. A model that produces slightly stronger answers may still be the wrong choice if the application cannot meet its service target under peak load.
Model size is only one variable. Deployment type and capacity matter too. Standard deployments are designed for variable demand and best-effort throughput, while provisioned capacity is intended for predictable performance at sustained volume. Global routing can offer broader capacity than geography-constrained options but changes where inference processing can occur.
This is why latency, throughput, and cost must be evaluated together. Model selection cannot be completed without a realistic load profile.
Cost should be evaluated per successful task
Token price is useful but incomplete. A cheaper model that needs repeated calls, larger prompts, more retries, or heavier retrieval context can cost more per successful user task. A more capable model may reduce orchestration complexity or allow the application to use shorter prompts. Batch deployment can dramatically change economics for delay-tolerant work.
Estimate unit economics at the workflow level. Include input tokens, output tokens, embedding calls, retrieval, safety checks, agent tool calls, retries, and any provisioned capacity. Then measure how often the workflow actually succeeds. The relevant number is cost per acceptable outcome, not merely cost per million tokens.
A good selection process can also use different models for different stages. A small model may handle routing or extraction while a stronger model is reserved for complex reasoning. That is often more efficient than forcing every request through the most capable model.
Reliability includes behavior under limits and failure
Production reliability is not only answer quality. The application must handle quota exhaustion, transient errors, timeouts, model unavailability, content-filter responses, malformed structured output, and long-running tool calls. Candidate models should be tested with the same retry, timeout, and fallback logic the production service will use.
Fallback should be intentional. Switching to a different model can alter output style, supported parameters, safety behavior, context limits, or tool-calling quality. A fallback path needs its own evaluation instead of assuming compatibility.
Model reliability is therefore part of model selection. The strongest candidate is one the platform team can operate, observe, and recover, not merely one that wins a static quality test.
Deployment compatibility can decide the winner
Foundry currently supports serverless API deployment for a wide range of models, with deployment types that include global, data-zone, geography-based standard, provisioned, and batch options. Some open-source, partner, or custom models can use managed compute. Not every model is available under every deployment type or region.
That means data residency and throughput requirements can shrink the candidate set dramatically. If the workload must keep inference processing within a particular geography, a globally routed option may be unacceptable even if it launches new models sooner. If the workload has predictable sustained volume, provisioned throughput may matter more than the cheapest pay-per-token rate.
The article on deployment models treats those differences as a first-class architecture decision. Model and deployment type should be selected together.
Use more than one model only when the routing logic is clear
Multi-model architectures can improve economics and quality, but routing introduces another system to test. The router must decide which requests deserve a larger model, which can use a smaller one, and what happens when the preferred model is unavailable. If routing quality is poor, the architecture can become harder to debug than a single-model design.
Start with clear rules or a narrow model-router use case. Measure task success by route. Keep model-specific prompts and parameters versioned. Do not let the router become an invisible source of behavior change.
Foundation-model choice also affects the user experience. Users experience the aggregate behavior of the routing, prompts, retrieval, and model—not the model card in isolation.
Choose with evidence, then keep re-evaluating
A practical process is straightforward: define hard constraints, create a representative evaluation set, shortlist models that are actually deployable, compare quality and safety, load-test realistic traffic, model the workflow cost, validate failure behavior, and document the decision. Keep at least one plausible alternative so the team has an upgrade or fallback path.
Model selection is not permanent. New versions appear, older versions retire, pricing changes, deployment types expand, and workloads evolve. Re-run the evaluation set before switching. A model upgrade should be treated like a software release, not a catalog refresh.
For teams working through Microsoft certifications, that discipline is more valuable than memorizing which model is currently strongest. The durable skill is knowing how to choose a model that fits the application’s real constraints and how to prove that choice with evidence.
One more useful discipline is to keep the model decision reversible. Store evaluation inputs outside the model provider, version prompts independently, and avoid coupling the application to provider-specific response fields unless the capability is genuinely required. Portability does not mean every model must be interchangeable; it means the team understands which assumptions would have to change if the selected model becomes unavailable, retires, or is surpassed by a better option. That makes model migration an engineering project with known dependencies instead of a rewrite discovered under deadline.