Microsoft AI-103: Visual Question Answering, Captions, and Alt Text
A multimodal assistant should describe and reason about what is visible, not invent details that sound plausible. In Microsoft AI-103, computer vision includes visual question answering, captions for one or multiple images, accessible alternative text, visual characteristics extracted with Content Understanding, and interpretation of image or video regions. The engineering challenge is to preserve source evidence while selecting the right output level for a user, a document pipeline or an assistive experience.
On this page
- Separate vision understanding from image generation
- Build a visual question-answering contract
- Write captions for different audiences and tasks
- Generate accurate alternative text and long descriptions
- Extract visual information using Content Understanding
- Work with video clips and temporal evidence
- Defend against visual prompt injection and unsafe media
- Evaluate and operate the vision pipeline
Separate vision understanding from image generation
Vision understanding accepts existing images or video segments and asks what information they contain. Image generation creates or edits media in response to a request. They may share a multimodal model family but have different acceptance tests: a generated image is judged against a creative brief, while an image description must be grounded in pixels that are actually present. Do not assume an expressive answer is an accurate one. A short caption might be sufficient for a decorative photo but inadequate for a screenshot showing a failure message.
Choose inputs deliberately. A technical diagram may need a high-resolution image or cropped regions so small labels remain readable; a set of chronological screenshots may need ordering context. If the service supports both image and text inputs, give an instruction about the task and the output format, but keep labels or instructions contained in the image in the lower-trust evidence layer. Document any preprocessing that changes the image before analysis.
Build a visual question-answering contract
A visual Q&A request should identify the question, authorized input images, and whether answers require a cited region or a concise uncertainty statement. Ask factual questions the image can support, such as how many warning icons appear or which status is displayed in a UI panel. Contrast those with questions the image cannot answer, such as why a person made a decision or whether a machine will fail tomorrow. An assistant should not invent a causal explanation from a single frame.
A grounded-answer rubric should require correctness, evidence localization when available, appropriate uncertainty and no unsupported references to hidden context. For a lab, use a small image collection with near-identical screens showing different error codes, and add a low-resolution copy. Evaluate whether the answer changes appropriately as evidence quality decreases. Some tasks need cropping or OCR as a preprocessing step; others require a model capable of interpreting the full visual relationship.
Where a question contains an object reference such as ‘the highlighted connector,’ the application must preserve which image region or annotation the phrase denotes. If the question and image arrive in separate requests, keep a stable asset identifier and protect the association with the user’s session. Otherwise an assistant can answer the right question about the wrong photograph. When the model’s evidence is ambiguous, a short clarifying question is preferable to silently selecting one of several similar objects.
Validation should also distinguish perception errors from reasoning errors. If the model reads ‘8.5’ from a chart whose actual label is ‘3.5,’ the initial failure is recognition; if it reads the correct value but calculates a trend incorrectly, the subsequent reasoning failed. Track those classes separately so the team can decide whether to improve image preprocessing, extraction, question framing or downstream calculation.
Use the following visual question-answering test: the input contains a line chart with three visible monthly values, but its legend is cropped out. The model may describe the plotted trends and cite visible axis labels; it must not invent which business unit each line represents. Give the evaluator three labels—supported, not visible, and ambiguous—and require the model to choose one for each assertion. This matters more than surface fluency when decisions depend on visual evidence.
Write captions for different audiences and tasks
A useful caption is contextual. An asset-management thumbnail may need object and setting; a technical screenshot may need the error text and relevant UI state; an analyst reading a chart may need the trend, axes and conspicuous exceptions. Requesting a generic description without specifying use can yield fluent but unhelpful prose. Decide whether captions should be short, detailed, structured or connected to bounding boxes or image segments based on downstream consumption.
For multiple images, the model can be asked to compare changes, but input order and identity must be explicit. If the application gives three screenshots, label their sequence and record which image supported each observation. Avoid writing a merged paragraph that confuses elements across separate frames. Test on mixed-format images and detect where a model transposes numbers or combines labels from different assets.
Generate accurate alternative text and long descriptions
Alternative text is not simply every visible object listed in order. It should convey the function of an image in its content context. A decorative flourish may require no substantive alt text; a chart representing enrollment trends may need a concise summary plus an accessible data table or extended description. AI suggestions can accelerate the first draft, but a reviewer should confirm the facts, particularly numbers, relationships, and content not visible at a glance.
Test generated alt text with a checklist: does it identify essential content, avoid redundant prefixes, omit speculation about protected attributes or intent, and respect the surrounding page? Where an image includes text, decide whether quoting the text is necessary for the user’s task. If essential chart values are too small to read, the answer should acknowledge the limitation rather than invent precise figures. The output must remain useful to someone who cannot see the image.
Consider an image showing a red emergency-stop button next to a clearly legible “PRESS TO STOP” label. A concise alternative text is “Red emergency-stop button labeled PRESS TO STOP on a gray control panel.” Do not add “the machine is unsafe” unless the surrounding document establishes that claim. A longer description can locate the button relative to adjacent controls when that position is necessary to complete a task; decorative color details should not crowd out the meaningful label.
Test the same image at lower resolution and with the label obscured. If OCR no longer supports the words, the system should say the inscription is unclear and flag the alt text for human review. Image-specific alternative text is an accessibility decision shaped by the page’s purpose, not a universal caption generated without context.
Extract visual information using Content Understanding
Microsoft’s blueprint includes configuring Azure Content Understanding in Foundry Tools to extract visual characteristics and video segments. Use the supported analyzer suited to the input modality and desired structured fields. The task may require reading both visible text and nontext visual features, such as a damaged component or diagram relationship. Do not confuse a raw OCR transcript with a complete semantic representation of an image.
Capture which part of the source supports each extracted value. For documents, a page number and bounding region may aid manual review; for video, timestamps and frames are critical. Where a model returns a structured object, validate it against a schema and business rules. If a field is absent or unclear, preserve an explicit missing status rather than filling a plausible default. In Content Understanding for documents, page numbers and bounding regions serve the same evidence function that timestamps serve in recorded media.
For structured visual extraction, review Microsoft’s Content Understanding analyzer configuration. Base analyzers such as prebuilt-image process image input, while custom field schemas define the values and types returned. Preserve the image ID, relevant region or other available evidence, and whether each field was directly visible. The current generally available multimodal service supports images, audio and video, but the older pro-mode API is retired; agentic mode is a separate document-only 2026-06-01-preview workflow, limited to one input file per request and no extract-method fields; it is not an image or video analyzer. Select a supported current image analyzer and validate its actual output instead of assuming a pro-image configuration exists.
Work with video clips and temporal evidence
Video introduces ordering and continuity. A model may need to process frames or segments to summarize events, but a final answer should preserve when the relevant evidence occurred. A video can contain an action at the beginning and an unrelated state at the end; treating both as simultaneous is a reasoning error. For long media, decide how clips are sampled, chunked and summarized without discarding the part needed to answer the user’s question.
Test a video where the scene changes before a key statement. Ask what changed and when, then compare generated answers with annotated timestamps. Include a negative test in which the requested event is not present. If a model detects objects, people or locations, apply task-appropriate privacy safeguards and avoid identity speculation. Multimodal understanding is not automatically forensic evidence or a substitute for a human specialist.
Defend against visual prompt injection and unsafe media
An image or screenshot may contain printed text telling an assistant to ignore its instructions or call a particular API. That text is data supplied by an untrusted source, not an authorization to act. Design the application so tool permissions, approval requirements and system rules are controlled outside the image. Test with a screenshot that mixes a legitimate error message with a malicious action request, and confirm that the system answers the user’s visual question without following the embedded command.
Content filters should classify prohibited material according to the application’s policy and platform features. Keep separate logs for an image that cannot be processed, content that is disallowed, and an answer that fails grounding evaluation. A refusal may protect a sensitive workflow, but overly broad rules can also prevent legitimate accessibility or safety use; document how such cases are reviewed by authorized people.
Evaluate and operate the vision pipeline
Create an evaluation set covering charts, OCR-heavy images, low-quality photos, multi-image comparisons, visual Q&A, alt text and short video segments. Review factual correctness and usefulness separately. A one-line caption can be faithful but omit necessary accessibility information; a long description can be detailed but invent a number. Include human-reviewable source references and record the evaluator rubric, model version, preprocessing and failed cases so results are reproducible.
The Microsoft AI-103 exam groups these tasks under computer vision, not under a generic image-generation shortcut. A practical checklist is task choice, input preparation, supported Foundry tool, evidence preservation, output schema, quality evaluation, safety policy, and a tested correction path. For audio and video content extraction, frame and timestamp identifiers provide reviewable support that static image descriptions alone cannot supply.
An audit-ready small dataset contains a product photo with correct text, a crowded scene with occlusion, a chart without a legend, and one adversarial screenshot containing a fake system instruction. For each expected answer, retain the visible supporting element and a clear reason to abstain when evidence is missing. Score descriptions for factual grounding and task utility separately from style. If the model obeys text in the image as an instruction, that is a prompt-injection failure even when the final caption looks normal.