Microsoft AI-103: Image and Video Generation, Editing, and Safety
The current Microsoft AI-103 computer-vision objectives explicitly include generating and editing images and videos, not only describing existing media. A production workflow must translate creative intent into an appropriate model request, manage reference assets and masks, apply supported generation controls, evaluate output quality, and reject content that violates policy or input rights. This guide separates generation, image editing, video creation and responsible multimodal review so that exam preparation does not confuse four different operational tasks.
On this page
- Choose the task and model before writing a prompt
- Generate images from text or reference media
- Use inpainting and mask-based edits appropriately
- Design a reliable video-generation workflow
- Edit generated videos without losing provenance
- Apply content safety and visual policy consistently
- Evaluate fidelity, usability and accessible output
- Prepare a bounded exam lab and operating checklist
Choose the task and model before writing a prompt
Begin with the intended artifact. A product mockup generated from a text description, a photo edited to remove a background object, and a short video created from a storyboard have different inputs, output formats, latency and approval needs. Confirm that the model and deployment available in your Foundry region support the requested media type and control features. Do not assume every multimodal model can both generate and edit images, or that video operations use the same synchronous response pattern as a short text completion.
Write down the requirements that determine success: subject identity, prohibited content, framing, aspect ratio, acceptable variation, accessibility, style consistency, rights to reference media and output use. Keep a human-readable provenance record of the chosen model, version, prompt, input assets and resulting artifact. This turns a subjective visual review into a traceable process when an image must be revised or an output challenged later.
Generate images from text or reference media
A generation prompt should specify the subject and the constraints that matter without trying to prescribe every insignificant pixel. If the request includes a supplied image, determine whether the service treats it as an image to transform, a style or content reference, or a starting point for an edit. Those capabilities and input sizes depend on the chosen model. A system should preserve user intent while communicating that stochastic generation may not reproduce the same composition twice.
For a lab, create three product illustration requests with a stable visual brief and a small permitted reference set. Vary one element at a time—such as environment, lighting or framing—and evaluate which changes the model actually follows. Reject assets with illegible text, distorted product features or invented logos. Compare the output against documented acceptance criteria, not merely whether it looks impressive on the first attempt.
A bounded image exercise starts with a task contract: “Generate a square product illustration against a neutral background; retain an exact supplied logo and avoid adding labels.” Test whether the chosen model actually supports reference images, the required aspect ratio and any fidelity controls. Preserve the model identifier and original image for human comparison; generative models can change typography or brand marks even when a prompt asks them not to. Reject rather than quietly deliver output that violates required text or identity constraints.
Use inpainting and mask-based edits appropriately
Editing should start with an explicit source image, the region allowed to change, and the attributes that must stay unchanged. A mask can identify pixels for replacement, but the model may alter boundaries, shadows or nearby features. Test whether an edited result still depicts the same authorized asset. If the source is an identity document, a safety-critical sign or an evidentiary photograph, stricter provenance and misuse controls may be necessary; do not represent a synthetic edit as an authentic record.
A useful negative test is to place a protected brand mark just outside the editable mask and request a background change. Review whether the mark is preserved and whether the output creates misleading annotations or visual artifacts. Keep the original image and mask, not only the final picture, so reviewers can reconstruct how the derivative was created. The exam measures configuring supported editing workflows, so inspect current model-specific instructions instead of inferring feature parity.
For an inpainting evaluation, supply a mask that covers only an unwanted background sign. Compare output outside the mask against the original image: an edit that removes the sign but changes an unmasked person’s clothing fails the preservation requirement. Record mask dimensions, image format and any service-specific rules before sending a request. A prompt saying “change only the sign” is not a substitute for checking whether the deployed image model offers the needed mask or edit operation.
Design a reliable video-generation workflow
Video generation is generally a larger job than producing a single image because it adds motion, timing, scene continuity and possible audio. Review the service’s supported inputs, output duration, asynchronous job semantics, quotas and content policy before accepting a user request. A production app may need to create a job, poll or receive status updates, store the resulting media with a controlled lifetime and recover gracefully when the job fails or is moderated.
A storyboard should describe a short sequence of actions and intended camera motion rather than a pile of unrelated adjectives. After generation, inspect continuity of faces, hands, objects, text, physics and scene order across frames. A clip that looks plausible in one screenshot may be unusable when played at normal speed. Keep validation distinct from the request status: a job marked complete can still produce an output that needs rejection or human editing.
A job queue becomes important when multiple users submit media-generation requests at once. Limit concurrency against quota, expose progress without guaranteeing precise completion time, and handle cancellation or expiration as real states. If a user refreshes the page, the application should reattach to an existing authorized job rather than submit a duplicate expensive request. Expired outputs should be removed under a documented retention rule, but provenance and review decisions can be retained separately when required for audit.
Inspect how a failure is surfaced when a media job is rejected by policy, exceeds a service limit or stops partway through processing. Those situations call for different user guidance. A moderated request should not be retried by slightly altering hidden content to bypass controls. A transient transport failure may permit a carefully bounded retry if the platform confirms that the original job was not already created.
Microsoft’s Sora 2 video generation documentation describes the model as preview in Foundry. In that documented preview, text-to-video and supported reference-media workflows run through an asynchronous job: submit an authorized request, retain the job ID, poll for succeeded/failed/cancelled status, and download the output only after a successful result. Retry with bounds and preserve cancellation/error details. The exact available model, supported region, time limits and deployment options must be checked before the lab, rather than assuming that every Azure subscription can run it.
For a concrete storyboarding test, request a five-second still-camera shot of a ceramic cup being placed on a desk. The acceptance sheet should record requested framing and duration, actual output resolution and length, whether hand-object contact is physically coherent, and whether moderation prevented a disallowed request. A failed or unavailable preview deployment is a documented lab outcome, not a reason to describe a video as successfully generated.
Edit generated videos without losing provenance
The AI-103 blueprint also mentions video-editing workflows. Treat this as a transformation chain: original permitted input, generated candidate, requested edit, edited output and approval. Determine whether the chosen platform supports prompt-driven editing, reference-conditioned editing, trimming or another operation. Do not promise a feature merely because a different model demonstrated it elsewhere. The system should preserve metadata needed to relate the final asset to the original request and its intervening changes.
For a lab, create a short generated clip and request one contained change, such as a different background or timing. Document which aspect is allowed to change and which aspects must remain stable. Compare the new clip at the same timestamps and inspect whether safety filters rerun on the modified version. Operationally, a second generation or edit can introduce fresh unwanted content, so approval is associated with each output revision.
The same Microsoft documentation describes a Sora 2 remix workflow that references a previously completed video ID and supplies a narrow revised prompt. Use an isolated test asset and change one property, such as the scene’s color palette; then compare it with the original for unexpected motion or scene changes. Do not equate prompt-based video remix with pixel-accurate frame masking, nor assume image inpainting controls automatically exist for video.
Video jobs and generated assets can have retention or expiration constraints. Store permitted outputs through your organization’s approved asset workflow, record rights/consent and provenance, and clean up temporary files. These procedures matter because an editing iteration without the source ID, model version and prompt cannot be evaluated or repeated.
Apply content safety and visual policy consistently
Generated media can create brand, privacy, impersonation and misleading-content risks. Configure the content safety features supported by the chosen model and apply organizational rules for acceptable subjects, reference-image permissions, watermark or provenance requirements and prohibited imagery. A watermark alone does not make an output safe, and a safety classifier is not a replacement for review when the use case is sensitive or the subject is an identifiable person.
Keep a policy decision log that distinguishes provider safety refusals from organization-specific rejection reasons. If a user tries to bypass a rule by embedding instructions inside a reference image, treat those pixels as untrusted task data, not as governing instructions. Extend prompt-injection defenses to image-derived text and other multimodal content before those inputs can influence privileged tools.
Evaluate fidelity, usability and accessible output
Use a review rubric that checks requested subject, scene consistency, recognizable objects, text legibility, temporal coherence and policy compliance. For product or training assets, include a reviewer who understands the real domain; visual realism alone cannot establish technical correctness. Record whether users can access a text alternative or caption describing the output. Generating a video and producing a meaningful description of it are separate activities requiring different validation.
For accessibility, create concise descriptive metadata for the final artifact and allow editors to correct errors. A decorative image may need a different alternative from a chart that carries essential data. A model may propose alt text, but it should not guess unseen identities or intentions. Accurate vision and multimodal AI separates what is visibly supported by an image from what the model merely infers, so editors can verify alt text before publication.
Ask a reviewer to compare both the creative result and the failure record. The rubric should separately score fidelity to the prompt, temporal consistency, visual artifacts, policy compliance and accessibility of any accompanying description. For a public-facing output, a “good-looking” generation still fails if it alters a protected logo, invents legible text that was not requested, or lacks evidence of required rights and approval.
Prepare a bounded exam lab and operating checklist
The Microsoft AI-103 study guide explicitly separates media generation/editing, multimodal understanding, and responsible visual AI. Build a small exercise for each instead of treating one text-to-image prompt as coverage of the whole computer-vision domain. Capture task choice, supported API, input preparation, generated output, edit parameters, safety result, review decision and cleanup. Record both successful and failed attempts.
Cost and artifact retention deserve attention. Media jobs can incur larger inference and storage costs than small text prompts; asynchronous failures can leave temporary assets or job records behind. Set limits on duration, request rate, saved history and who can retrieve finished files. Exam readiness means knowing how to choose the supported workflow and verify that the resulting asset is useful, lawful, traceable and safe.