Microsoft AB-100: Model Routing, AI TCO, and Build-or-Buy Choices
A lower model price does not guarantee a cheaper business process. An agent may retry a weak answer, call a tool unnecessarily, or route a sensitive task to the wrong deployment. AB-100 solution architects must compare options using business results, quality, operational cost, and control—not only token tariffs. Model routing can improve cost and latency by selecting a suitable model for each task, but it introduces its own quality and governance decisions. This guide builds a practical decision method for model choice and for the larger question of buying, extending, or building AI components.
On this page
Establish the unit of business value
Begin with a workflow whose output is measurable: accurately categorized service cases, approved sales summaries, resolved inventory exceptions, or safe contract reviews. Record baseline minutes per completed task, error rate, escalation effort, and rework. Separately record the consequences of a policy breach, incorrect recommendation, or delayed response. A process that is cheap per token but requires constant review may be more expensive than the manual baseline.
Use a denominator that reflects completed, acceptable outcomes. For example, cost per policy-compliant resolved case captures retries, escalations, supervision, and failed attempts, while cost per model invocation cannot. Build the economic model using observed pilot data, not an assumed percentage productivity gain. The business value article explains why adoption counts alone are not a return on investment.
Compare prebuilt, extended, and custom approaches
A prebuilt business application capability can shorten time to value and reduce maintenance, but it may not expose the exact controls a regulated workflow requires. Extending Microsoft 365 Copilot or Copilot Studio can reuse enterprise identity and channels while adding scoped actions or grounding. A code-first Microsoft Foundry solution can support specialized orchestration and custom model evaluation, at the cost of increased deployment responsibility.
The decision record should include user population, channel, required data residency, integration ownership, testing burden, license terms, service limits, fallback behavior, and upgrade path. Do not justify a custom stack merely because its first prototype looks impressive. Conversely, a low-code path that cannot enforce an approval gate may be inappropriate for a consequential write. See Copilot versus custom agents for a product-surface comparison.
Decide which requests are eligible for routing
A simple extractive classification question, a policy-grounded support answer, and a high-risk financial recommendation have different quality requirements. Assign each to a task class with an evaluation set, latency budget, cost budget, and required controls. Route only among models that satisfy all nonnegotiable constraints. A sensitive workload must not drift to a region or provider that violates the organization’s policies merely because routing predicts a better price.
Microsoft Foundry’s model router can select underlying models in supported configurations. That does not absolve the architect of checking model availability, tool behavior, deployment mode, output quality, latency, and billing rules. Design fallback behavior for a model that is unavailable or consistently misses a quality threshold, and record the selected model when telemetry permits.
Calculate total cost of ownership rather than API spend
Separate costs into design, licensing, data preparation, development, evaluation, inference, connectors, orchestration, observability, review labor, ongoing compliance, and model-change testing. Add a contingency for new document formats or changing business policies. Compare those costs against measurable savings such as reduced handling time or fewer escalations, not against an arbitrary target for prompts per employee.
A financially credible pilot also reports uncertainty. If case resolution improves during a busy season, ask whether traffic mix changed. Compare similar cohorts or use a staged rollout. The architect must be able to explain the sensitivity of ROI to failure rate, adoption, and maintenance overhead instead of publishing one optimistic annual saving estimate.
An illustrative cost worksheet should separate a base request, extra retrieval queries, fallback model calls, human review, API connector charges, logging, and incident handling. For each path, multiply observed frequency by the actual unit cost and add fixed operational costs; then divide by accepted business outcomes. If a cheaper route creates more rework, model price savings can disappear. Sensitivity tests should show what happens when adoption doubles or a model’s refusal rate increases.
Build an evaluation matrix for routing decisions
Construct a set of representative tasks and score each candidate on correctness, groundedness, unauthorized disclosure, tool success, latency, and total cost. Include cases requiring refusal, clarification, and human escalation. When prompts or tools change, re-run the matrix; an apparent cost saving can hide a new class of wrong answers. A router should be rolled out gradually with a rollback option and traceability to the model selected for a decision.
Make the matrix a release gate: if the low-cost route fails an access-control test, price cannot compensate. If a larger model performs no better on a routine classification task, spending more may be unjustified. This discipline ties architecture to enterprise outcomes rather than to allegiance to any one model family.
A quality gate should use the same task set for each eligible model and report results by class, not just as a single average. For a low-risk summary, relevant measures may include factual consistency and latency. For a high-impact account change, unauthorized disclosure or tool misuse should be an automatic rejection. Record which model handled each request for diagnosis while avoiding sensitive prompt leakage into telemetry. Model routing should be explainable as a policy decision even when selection uses learned performance signals.
Keep governance visible in the final recommendation
The business sponsor should receive a recommendation that names the selected design, the alternatives ruled out, cost range, major assumptions, security owner, service owner, and next review date. The operational team should know how to disable routing or force a safer fallback without redeploying the whole business system.
AB-100 preparation is about making such decisions coherently. The answer is rarely “always choose the strongest model” or “always choose the cheapest model.” It is to make the boundary conditions explicit, measure real task outcomes, and operate the architecture as a controlled business service.