Microsoft AB-100: End-to-End Tests Across Dynamics 365 Agents
An agent can pass a conversation test and still fail the business process it was built to support. A customer-service agent might summarize a case correctly while updating the wrong contact; a finance assistant might recommend an action that is valid in one legal entity but forbidden in another. AB-100 deployment objectives include end-to-end testing across Dynamics 365 applications. The architect must define what a correct system outcome looks like, then collect evidence across identity, channels, retrieval, tools, transaction logs, and human handoffs. This guide starts with system state rather than polished chat transcripts.
On this page
- Map the full business transaction across systems
- Construct a scenario matrix with expected outcomes
- Make tool and connector contracts observable
- Test grounding and user-visible explanations separately
- Include channels, people, and recovery in the test boundary
- Use the test results as a formal deployment gate
Map the full business transaction across systems
Use a scenario that crosses Sales, Customer Service, and Finance: a customer requests a change to an order while a support case is active. Identify the authoritative system for customer identity, credit status, order state, and customer communication. Draw data movement arrows and mark each required approval. The test should end with the correct persisted records and customer-visible confirmation, not merely an agent response saying that work is complete.
Separate synchronous decisions from asynchronous propagation. A CRM update may appear immediately while a finance posting or downstream notification occurs later. Tests need explicit eventual-consistency allowances and a deadline after which the workflow is considered failed. Without this contract, the test suite can either hide real faults or flag healthy workflows as broken.
Construct a scenario matrix with expected outcomes
Include routine read-only questions, low-risk approved writes, high-risk requests requiring human review, denied access, missing data, and stale or contradictory grounding. For each, specify the starting state, actor identity, input, expected visible response, permitted tool calls, forbidden actions, and final data state. Include one test where the user changes their mind after an agent has started planning but before a transaction is committed.
A strong matrix exercises both a successful and unsuccessful path through each critical interface. An agent that refuses an unauthorized request and creates no downstream record should count as a correct result. The human-oversight article helps define what should cross the approval boundary.
A useful matrix has at least four dimensions: the initiating role, the source system, the authorized downstream effect, and the failure state. A sales user can read an account note but may not see finance-only credit data; a service user may update a case but cannot approve a supplier payment. Create positive, denied, partial-failure, and recovery rows. Test setup data must be stable enough that a later failure can be attributed to the agent or integration rather than a moving business record.
Make tool and connector contracts observable
A business action needs typed arguments, validation rules, timeouts, clear error codes, and trace identifiers. Record whether an action was attempted, committed, declined, or timed out with uncertain result. Where appropriate, add an idempotency key and query the system of record before retrying a failed confirmation. This prevents a model or orchestration layer from converting a network blip into two orders or two refunds.
The test harness should correlate the initiating user with the tool execution principal. Verify that a person denied access in Dynamics 365 cannot gain it through a Copilot Studio connector or service account. Test row-level and legal-entity restrictions, not only broad role names. A correct API response code is insufficient if the returned data violates its access contract.
Test grounding and user-visible explanations separately
Answer quality and data integrity are different checks. Evaluate whether the agent cites the correct current policy and whether its reasoning uses facts the user was entitled to see. Then separately verify the executed action's stored result. A generated explanation may be well written while relying on a superseded policy; a correct business update can still be followed by a misleading success message.
Use documents with controlled version histories and deliberately revoke one source. The agent should not continue citing it after the allowed refresh window. Include prompt-injection text in a retrieved record and prove that it cannot override the approved tool contract. Record all failures with the same labels used by production monitoring so test findings inform alert design.
Include channels, people, and recovery in the test boundary
Run representative cases through chat, service-channel handoff, and embedded application interactions as supported by the environment. Test whether a human receives the case number, verified context, actions attempted, and reason for escalation. Verify how the application behaves when the agent is offline or a downstream model endpoint fails. The basic business application should continue to protect and save valid work.
Load testing should respect service limits, and performance measures should reflect the entire journey. A two-second answer may be irrelevant if the associated task remains incomplete for twenty minutes. Capture percentile completion time, rework, denied actions, duplicate attempts, handoff rate, and customer effort.
Use the test results as a formal deployment gate
A release report should show which scenarios ran, the source-data snapshot, deployed model and prompt versions, environment, permission set, known limitations, and approval owners. Require zero unmitigated critical permission or data-integrity defects before launch. Lower-severity language variations may be handled differently, but that decision should be deliberate and traceable.
Post-release telemetry must check the same conditions that the tests asserted. If production starts exhibiting more repeated tool calls or missing handoff context, the release team needs a way to compare that behavior against the approved baseline. The result is a controlled agentic business service, not merely an attractive demonstration.
A green response-quality score does not excuse a failed authorization or duplicate-write test. Define hard fail conditions for access-control escapes, missing approval, cross-tenant leakage, untraceable financial writes, and inconsistent transaction state. Other findings can be risk-scored by frequency and severity. A release report should identify the actual deployment configuration tested—including connector and model versions—so the team does not approve one configuration and silently publish another.