Practice Exams:

Microsoft AI-103: Document Extraction with OCR, Layout, and Field Validation

AI & Machine Learning

Documents contain structure that ordinary text extraction can destroy. A scanned invoice may have two similar totals, a policy can refer to a table on the next page, and a contract can use headings and footnotes to qualify an obligation. Microsoft AI-103 assesses information extraction using OCR, layout analysis, structured fields and Content Understanding in Foundry Tools. The goal is not only to read the text, but to create a trustworthy, reviewable representation that can support search, agents and business workflows.

On this page
  1. Choose a document representation from the business task
  2. Combine OCR with layout-aware extraction
  3. Design a schema for meaningful fields
  4. Use Content Understanding and markdown appropriately
  5. Prepare index chunks without discarding context
  6. Integrate custom enrichment and exception handling
  7. Protect sensitive documents and validate downstream actions
  8. Validate an end-to-end AI-103 extraction lab

Choose a document representation from the business task

Start by asking what the application needs: a searchable paragraph, a small set of invoice fields, a faithful markdown representation, table rows or an audit-ready source excerpt. Raw OCR may be sufficient for searching simple scans but insufficient when table columns, signatures or page ordering affect interpretation. A model-generated summary can be useful for browsing yet cannot replace evidence when a regulatory reviewer must verify the exact original clause. Define the output and its review requirements before picking an extraction model.

Document sources also vary. Digital PDFs, scans, photographs, spreadsheets and mixed slide decks may require different preprocessing and permissions. Record file version, page count, MIME type, language, owner and sensitivity. Never assume that because one parser handled a clean English invoice it can interpret all handwritten forms or complex cross-page tables.

Combine OCR with layout-aware extraction

OCR detects written text, but its output order may not match the document’s logical reading order. A two-column brochure can interleave paragraphs, and a financial table can lose the relationship between headers and numeric cells. Layout-aware analysis should help preserve sections, paragraphs, tables and image context where supported. Inspect the extracted representation on representative examples before feeding it to a generator or search index. An apparently complete text output can still be semantically damaged.

For a lab, use a clean digital contract, a low-resolution scanned invoice and a report with a table spanning two pages. Compare plain OCR text with a layout-aware output and identify missing units, merged columns, figure labels and separated row groups. Preserve page and region references where the service provides them. A downstream assistant should be able to show the exact source of a claimed value instead of pointing only to the full file.

A two-column purchase-order PDF illustrates why raw OCR text is not enough. If the left column lists quantities while the right column lists unit prices, flattened reading order can associate a value with the wrong item. Review table row geometry, multi-page headers and repeated footers before using the values in an automated workflow. The acceptance record should include both the extracted field and the source page or region where it was observed.

Design a schema for meaningful fields

Field extraction should distinguish amounts, currencies, dates, account identifiers, line items and optional text with explicit data types and validation rules. An invoice might include subtotal, tax and total; choosing the largest monetary number is not a valid extraction method. Define how the pipeline handles multiple candidate values, missing information and unclear handwriting. A required field that cannot be read should trigger review, not a plausible invented value.

If the document contains repeated records, use arrays with stable item boundaries instead of concatenating all values into a single paragraph. Validate field types, dates, arithmetic consistency and allowed values after the model returns a candidate structure. For some tasks, a simple deterministic postprocessing rule catches errors that a sophisticated model leaves in place. Keep the source evidence associated with each critical field so reviewers can correct it efficiently.

Microsoft’s Content Understanding analyzer configuration describes fieldSchema for structured extraction. This compact illustrative analyzer configuration fragment requests two invoice fields. A real deployment must supply its current required analyzer properties, supported model and API version:

{
  "baseAnalyzerId": "prebuilt-document",
  "fieldSchema": {
    "fields": {
      "VendorName": {
        "type": "string",
        "description": "Supplier on the invoice",
        "method": "extract"
      },
      "InvoiceTotal": {
        "type": "number",
        "description": "Final amount due",
        "method": "extract"
      }
    }
  }
}

On an invoice with subtotal 900, tax 90 and grand total 990, the expected InvoiceTotal is 990 only if the source labels make it the amount due. A model that returns 900 may satisfy the JSON schema yet fail the business validation. Preserve the page and region supporting the total, compare with any available confidence/grounding signal, and route ambiguous figures for review rather than fabricating evidence coordinates.

Use Content Understanding and markdown appropriately

Microsoft Foundry’s Content Understanding tools can extract structured or markdown-like representations suited to downstream reasoning in supported scenarios. Markdown can retain headings, lists and table structure more effectively than a flat string, but it is still a derived representation, not the original document. It may omit or reinterpret an image or table depending on configuration. Validate whether the chosen analyzer and API version support the required modality, region and outputs using current Microsoft documentation.

Choose configuration according to the application: a concise structured schema for transactional automation, faithful layout content for search and summarization, or combined outputs where both are needed. Keep the model/version and analyzer configuration with the extracted artifact. When processing changes, re-run a representative evaluation set to see whether table recognition, field completeness or evidence references improved or regressed.

For a current document-extraction pipeline, start with generally available 2025-11-01 analyzers such as prebuilt-read and prebuilt-layout, then choose a field schema and source-grounding settings appropriate to the document. Confidence and grounding are configured on supported document analyzers; do not assume every response or modality supplies field confidence. The legacy pro-mode API 2025-05-01-preview is retired. Agentic document analysis instead offers preview reasoning in 2026-06-01-preview, but supports one input file per request and does not support schema fields using the extract method. Decide whether GA extraction or preview agentic reasoning matches the evidence task; do not copy an older pro-mode analyzer configuration or claim they are interchangeable.

Prepare index chunks without discarding context

Retrieval systems work best with chunks that preserve meaningful semantic units and source metadata. A policy section may be one unit, while a long manual may need subdivision by heading or topic. A table might require keeping its header with the relevant rows; splitting that header away can make a correct retrieved passage misleading. Use page, section, document version, access scope and last-update metadata to support filtering and accurate citations.

Compare keyword, vector and hybrid retrieval on questions that require exact codes and conceptual language. A vector hit that retrieves a similar policy from the wrong version is a retrieval failure, not a reason for the language model to improvise. After parsing, Azure AI Search retrieval for RAG must retain the source version, page number and access-control metadata so a plausible hit from an old policy is not mistaken for current evidence.

Integrate custom enrichment and exception handling

A retrieval pipeline may enrich extracted content with labels, normalized units, image descriptions or derived summaries. Treat each enrichment as a separate transformation that can fail or change facts. If an image is described incorrectly, adding its description to a search index can make the wrong interpretation more discoverable. Review important enrichments against originals, log the transformation type and version, and distinguish the source’s assertions from the model’s commentary.

Route unsupported formats, corrupt scans and ambiguous fields to an explicit exception workflow. A partly successful ingestion job should not mark a document ready if its critical pages or fields are missing. Record which pages were processed, which failed, and whether the file is searchable but not approved for automated action. This prevents a downstream agent from mistaking an incomplete index for a complete record of business evidence.

Ingestion versioning matters whenever a document changes after it was indexed. Record a stable document ID, file version or content hash, extraction version and last successful indexing time. When a corrected document arrives, replace the old indexed representation deliberately rather than leaving two near-identical versions discoverable without a freshness rule. A model that cites a superseded policy can mislead users even when retrieval and generation were functioning exactly as configured.

Add reconciliation between the source repository and the search index. A failed deletion or expired permission can leave an index containing data the user should no longer retrieve. Test that updating or removing a document also updates its chunks and access metadata, and alert on ingestion failures that persist beyond the freshness target. The consistency contract is part of trustworthy information extraction, not just search administration.

Protect sensitive documents and validate downstream actions

Documents often contain customer addresses, signatures, health details or confidential financial records. Classify them and apply least-privilege access to storage, index and extraction results. Security trimming should follow the requesting user’s authorization even when the embedding or text index contains a representation of the content. A model must not reveal fields from a document simply because it can retrieve a relevant vector. Define retention and deletion rules for raw files, temporary extraction assets, logs and indexed derivatives.

Treat document content as untrusted data. A scanned page might contain instructions telling an AI agent to call a tool or ignore access controls; extracting that text does not grant it authority. A workflow proposing a payment or account update must validate the extracted fields and require business approval independently. Effective prompt-injection defenses prevent extracted page text from authorizing tools or changing application-level instructions.

Validate an end-to-end AI-103 extraction lab

Build a small pipeline with three controlled documents: a digital contract, a scanned invoice and a multi-page report. Extract layout and fields, validate schemas, index useful sections, then ask an agent questions requiring exact evidence. Check a negative question whose answer is absent and verify the system abstains. Measure field accuracy, table reconstruction, citation usefulness, ingestion time and cost. A successful document upload is not proof that answers will be grounded in correct structure.

The Microsoft AI-103 exam includes content extraction alongside multimodal ingestion and search. Review the specific official objectives before treating one OCR experiment as complete coverage. Content Understanding for complex documents exposes layout-dependent ambiguity in tables, forms and low-quality scans; audio and video content extraction instead requires segment and timestamp evidence.

A useful negative test has an invoice with a duplicated total, a missing supplier name and a stamped approval note placed across a line-item table. Keep expected values and document evidence in the fixture. The review succeeds only when the pipeline identifies unresolved fields, retains the correct page/region, and prevents extracted text from instructing downstream tools. Both false certainty and silent document omission count as failures.

Continue learning

Related guides

Content Understanding for Messy DocumentsBusiness documents rarely arrive as clean paragraphs ready for a language model.RAG on Azure: Retrieval Quality Is the ProductRetrieval-augmented generation is often presented as a simple pipeline: embed documents, search for the nearest chunks, place them in a prompt, and let a language model answer.Azure AI Search for RAGAzure AI Search can serve as the retrieval layer for a RAG application, but its value is not simply that it stores vectors.Grounding Azure AI with Enterprise DataEnterprise grounding is not simply connecting a model to documents.

Related Posts

• Mastering the Basics: Your Guide to Microsoft Azure AI

• Microsoft AI-300: How to Roll Back a Model Without Rolling Back the App

• Microsoft AI-103: Managed Identity for AI Application Credentials

• Microsoft AI-901: Computer Vision, Language, and Generative AI

• Microsoft AI-103: Agent Identity in Azure AI Foundry

• Microsoft AI-103: Cost Control for Azure AI Apps

• Microsoft AI-103: Latency Tuning for Azure AI Apps

• Microsoft AI-103: Reproducible ML Pipelines on Azure

• Microsoft AB-100: Copilot Adoption Without AI Sprawl

• Microsoft AB-100: GitHub Copilot Repository Instructions