Microsoft AI-103: Text Analytics, Translation, and Structured JSON
Text analysis in Microsoft AI-103 connects language understanding to usable software outputs. The exam includes extracting entities, topics and summaries; recognizing sentiment, tone and sensitive content; translating languages; and generating structured results for domain-specific tasks. The engineering challenge is not simply asking a model for JSON. It is defining a measurable contract, distinguishing a missing fact from an incorrect one, choosing supported Foundry or Translator capabilities, and validating the result before a downstream workflow trusts it.
On this page
- Select the task and the evidence it requires
- Extract entities and topics with clear definitions
- Produce structured JSON that downstream code can validate
- Summarize without introducing unsupported claims
- Classify sentiment, tone and sensitive material appropriately
- Translate text with terminology and source context
- Integrate text analysis into agent workflows safely
- Demonstrate AI-103 readiness with failure-focused labs
Select the task and the evidence it requires
An entity extraction service needs a schema of fields and types; topic classification needs an approved category set; a summary needs a clearly specified audience and evidence boundary. A customer email may ask about one invoice and include several reference numbers. If an application extracts the wrong identifier, downstream automation can act on a different customer even when the summary is persuasive. Start by defining input sources, authorized data use, permitted outputs and how unknown values must be represented.
Model choice follows those needs. A compact model may classify a limited taxonomy effectively, while a richer language model may be needed for a multi-document compliance summary. Foundry Tools and Azure Translator may also fit specialized steps. Compare latency, cost, supported languages and validation capability with a small representative dataset before selecting a service. The goal is not to find the most impressive demo; it is to return the right fields for real cases.
Extract entities and topics with clear definitions
Entities should be tied to useful domain categories such as product ID, monetary amount, organization, expiration date or case reference. The application must decide whether two surface forms refer to the same entity and what to do when dates or numbers conflict. Extracting every proper noun from a document may create clutter rather than a useful data model. Use a schema that includes source span or evidence reference for fields that require review, and avoid silently inferring values absent from the text.
Topic classification presents a separate problem. One support request can concern billing and account security simultaneously, so a single-label classifier may lose information. Decide whether the task needs one category, multiple labels or a hierarchical taxonomy, and provide examples of edge cases in evaluation. Check precision and recall for categories that carry operational risk. Make the model abstain or route to review when the evidence does not satisfy its classification criteria.
Produce structured JSON that downstream code can validate
A JSON response should have a defined shape, required fields, allowed value ranges and error semantics. Where the model or API supports schema-constrained outputs, use them while still applying business validation. Syntactically valid JSON can contain an impossible date, a currency with the wrong units or an account number belonging to another document. Validation needs to be layered: parse the result, check the schema, compare critical fields with source evidence, and reject or review contradictions.
For a lab, provide three invoices with missing tax IDs, multiple totals and mixed currencies. Extract invoice date, total, vendor and payment terms; require null or a clear missing-state for absent values. Deliberately ask the system to fill a nonexistent field and observe whether it invents one. Log mismatches without storing full sensitive records in a broad analytics workspace. The useful skill is reliable integration, not merely getting a model to print curly braces.
Schema versions are part of the service contract. If the application adds a new required field, older documents may legitimately lack that information. Plan a migration or explicitly support two schema versions rather than treating every missing value as a model failure. Store the extraction schema version with each result, log rejected fields and maintain a review path for records that cannot be converted safely. This is especially important when agents share outputs across services owned by different teams.
Measure validity at multiple levels: percentage of responses that parse, percentage that satisfy types and required fields, percentage whose claims match source evidence, and percentage the business can accept without manual correction. A model can achieve near-perfect JSON syntax while frequently misidentifying invoice dates. Operational dashboards should expose that distinction so teams do not optimize only the easy metric.
Suppose a customer message says, “My C-204 access was restored, but I was charged twice and need a refund.” Define separate requirements for the case identifier, issue labels, refund intent and a short evidence quote. A useful application-level result is shown below. This is an illustrative business schema, not a verbatim Azure service response:
{
"case_id": "C-204",
"issues": ["billing", "access"],
"refund_requested": true,
"evidence": "charged twice and need a refund"
}Validate that case_id matches the permitted identifier pattern, issues comes from an approved enum, the refund flag is Boolean, and the quoted evidence exists in the source. Send the same extractor an item saying only “please review my bill”; it must not fabricate a case ID or force refund_requested=true. Schema validity, source-supported facts and permission to act on the result are three separate tests.
Summarize without introducing unsupported claims
Summaries are selective representations, so choose what must be retained before drafting. A compliance summary may need obligations, exceptions, effective dates and supporting clauses; a customer-support summary may need the problem, actions already attempted and unresolved next steps. Do not ask for a confident short answer when conflicting records have not been reconciled. Preserve references to the original document or conversation turns when readers must verify the condensed claim.
Evaluate a summary against a reference set that includes critical facts, unsupported assertions and omission risk. Two summaries may sound equally fluent while one drops a deadline or changes the order of a decision. Test long input, contradictory paragraphs and hidden instructions embedded in source material. Treat source text as evidence rather than a command channel, and verify that the summarizer does not follow prompts contained in a customer message.
Classify sentiment, tone and sensitive material appropriately
Sentiment can vary within a single message. A customer may praise the technical support while criticizing billing; a single positive label would hide the actual service issue. Tone analysis is similarly context-dependent and affected by sarcasm, cultural expectations and dialect. Use these signals to support review and prioritization, not as authoritative diagnoses of an individual’s character or emotional state. Assess error patterns across the languages and content types the system will actually see.
Sensitive-content detection needs a defined policy, not just a classifier threshold. Decide what categories are prohibited, which are allowed with warnings and what should trigger human review. Measure false positives and negatives, especially where the application serves legitimate discussions of difficult subjects. Use Azure AI Content Safety to support category-specific policy checks, while reviewing legitimate borderline cases rather than treating every classifier flag as a final verdict.
Translate text with terminology and source context
A translation service should preserve meaning, proper nouns, product codes, legal phrasing and units even when natural sentence structure differs across languages. A general LLM may paraphrase too freely for a regulated document; a domain translation workflow may require terminology guidance and bilingual review. Check whether the chosen Azure Translator or Foundry-supported path has suitable language coverage and API behavior. Never assume that because two models accept the same languages, they offer the same fidelity or terminology controls.
Build a test set containing ordinary text, technical instructions, a contract clause, dates, numbers and ambiguous acronyms. Compare machine translation with vetted reference versions, mark unacceptable meaning changes and capture revision reasons. When language detection is uncertain, ask for clarification or surface the uncertainty. For live interaction, evaluate latency and how repeated short segments affect sentence coherence. Microsoft’s Translator documentation defines current service capabilities and should be checked when API versions change.
Microsoft’s Translator 2026-06-06 migration guidance explicitly warns that the generally available API is not a drop-in replacement for v3.0. The new REST contract uses inputs for the request collection, targets for requested languages and value in the response; older parsers relying on the v3.0 response shape can silently corrupt a downstream workflow if the change is ignored. A small request body from the current translate API has this structure:
{
"inputs": [{
"text": "Please confirm the appointment.",
"language": "en",
"targets": [{"language": "es"}]
}]
}When designing a migration, pin the chosen API version, validate the response against that version, and test names, amounts, negation and protected technical terms. Authentication and endpoint choice depend on the configured Azure resource; do not assume the new request can be posted to an older v3.0 path unchanged.
Integrate text analysis into agent workflows safely
An agent may use extracted entities as arguments to another service, but that is an authorization boundary. A vendor name found in untrusted text is not permission to execute payment or account changes. Validate the schema and authenticated user’s authority before dispatching a tool call, and require approval for significant actions. If an extracted number is uncertain, the workflow should request confirmation rather than silently defaulting to the nearest plausible value.
Monitor each stage separately: input decoding, language detection, extraction, validation, translation, model reasoning and tool use. Failures have different remedies. An incorrect field may indicate schema ambiguity or missing source evidence; an error response may indicate throttling or expired credentials; a tool denial may be exactly the expected security control. End-to-end tracing helps avoid fixing one problem by expanding permissions or hiding validation exceptions.
Demonstrate AI-103 readiness with failure-focused labs
The current Microsoft AI-103 skills outline tests entity and topic extraction, summaries, structured outputs, sentiment/tone/safety and text translation. Prepare a small lab for each task and document both a good result and a known failure. For audio input, speech-to-text and translation add recognition and language-selection uncertainty before the text-analysis stage begins.
Review output evidence, model/service choice, data handling, error paths and operational metrics. A useful test of readiness is whether you can explain why the classifier or translator was chosen, how critical output is checked, and what happens when a required field is missing. If those answers are clear, the solution is much closer to a working production integration than a prompt-only prototype.
An effective cross-task failure set contains an ambiguous entity, an invalid enum, an answer with unsupported sentiment, a translation with reversed negation and an instruction embedded in an extracted quote. Track each error by stage—input, extraction, JSON validation, translation or agent tool authorization. If a model produces a valid JSON structure with a wrong account number, parsing succeeded but the business result still failed.