Practice Exams:

Computer Vision, Language, and Generative AI: Start With the Problem

 

Teams often choose an AI service before they have defined the problem. A stakeholder asks for “AI,” someone proposes a chatbot, and only later does the group discover that the real task was to extract fields from documents, classify images, detect entities in text, or generate a draft from approved source material. Better AI design starts by framing the information transformation before selecting a model.

The current AI-901 exam makes that skill explicit. Candidates must identify common AI workloads and then implement lightweight solutions using Microsoft Foundry. That combination rewards people who can distinguish the problem first and the platform mechanism second.

For Azure AI Fundamentals, computer vision, text and speech, information extraction, generative AI, and agentic AI should therefore be treated as problem frames. Modern multimodal models may span several categories, but the business requirement still determines what success looks like.

Begin with the input and the decision the system must support

A useful problem statement identifies the input, the required output, and what decision follows. “Use AI on invoices” is vague. “Extract supplier, total, due date, and purchase-order number from incoming invoices so finance can validate and route them” is actionable. The latter tells the team it needs information extraction with defined fields and validation, not an open-ended conversational interface.

The same discipline applies to images, speech, and text. If the business needs a probability or category, predictive modeling may fit. If it needs structured facts from unstructured material, extraction may fit. If it needs new language or imagery, a generative model may fit. The solution becomes easier to evaluate once the transformation is specific.

A precise problem frame also identifies unacceptable outcomes. If an invoice field is missing, can a human correct it before payment? If an image classifier misses a safety defect, what is the consequence? If a generated response invents a product policy, can the answer reach a customer unchecked? Stating these failure conditions early helps determine whether the workflow needs confidence thresholds, review queues, deterministic validation, or a different technical approach.

It is also useful to state what the system should not decide. A tool can extract or summarize evidence without being authorized to approve a loan, diagnose a patient, or release a payment. Separating AI assistance from business authority keeps the problem frame precise and makes it easier to place human judgment at the points where consequence is highest.

Computer vision is about interpreting or producing visual information

Computer vision includes tasks such as image classification, object detection, visual inspection, optical character recognition, image analysis, and—in the generative era—visual question answering and image generation. A broad resource on computer vision helps show why these are not one problem. Detecting a defect on a production line is different from describing a photo or creating a marketing image.

The appropriate evaluation also changes. Object detection may require location accuracy and missed-object analysis. Document extraction may care about field-level correctness. Image generation may be judged on prompt adherence, brand constraints, safety, and visual quality. Calling all of these “vision AI” is correct but not sufficient for architecture.

Language workloads range from analysis to generation

Natural-language systems can detect entities, extract keywords, classify intent, analyze sentiment, summarize, translate, answer questions, or generate new text. A practical introduction to natural language understanding is useful for separating tasks that interpret language from tasks that create it.

This distinction matters because deterministic structure may be more valuable than fluency. If a compliance workflow needs to identify a contract party and expiration date, a constrained extraction process may be more appropriate than asking a general chat model to write a paragraph about the document. The best interface is the one that supports the business decision reliably.

Language solutions also differ in how much freedom the output should have. A sentiment score has a constrained result space; a summary is more flexible; a customer-facing response may need both grounding and policy constraints. Teams should resist using the same evaluation and governance pattern for all three merely because text is the input and output medium.

Generative AI is strongest when synthesis is part of the requirement

Generative AI makes sense when the desired output is new content: a draft, explanation, image, code suggestion, conversational response, or transformed version of existing material. It is particularly useful when there are many acceptable outputs and the system needs flexibility rather than a single deterministic mapping.

Understanding the difference between generative AI and large language models prevents another common framing error. An LLM is not the answer to every generative task, and a generative model is not automatically the best tool for a problem that only needs classification or field extraction.

Multimodal models combine capabilities but do not remove architecture choices

A multimodal model can accept text and images, sometimes audio or video, and produce language or other outputs. This enables elegant experiences: a user can photograph equipment and ask a question, upload a chart and request an explanation, or combine a document with instructions. Yet a single model endpoint does not eliminate the need to decide what data is allowed, how outputs are validated, and whether a specialized component would be more reliable.

Multimodality should be viewed as a capability expansion, not permission to skip problem decomposition. A production workflow may use one model for flexible reasoning while still relying on structured extraction, rules, search, or deterministic business logic for high-confidence steps.

Multimodal applications can also hide data-flow complexity. A user may upload an image that contains text, faces, account numbers, or location information even when the application is marketed as a visual assistant. The architecture should treat each input modality as a data source with its own sensitivity and retention implications rather than assuming an image is less sensitive than a text prompt.

Information extraction is different from free-form generation

Information extraction turns unstructured content into defined facts. The source might be a form, image, audio recording, or video. The desired output is usually structured and can often be validated against types, ranges, or business rules. That makes extraction a strong fit for automation workflows where downstream systems expect predictable fields.

Generative output is looser. A model can explain or summarize a document, but a downstream payment process should not depend on an unconstrained paragraph when it really needs an exact amount and account identifier. Separating extraction from generation improves both reliability and auditability.

Agents are appropriate when the problem includes a sequence of actions

An agent becomes relevant when the system must pursue a goal across multiple steps and tools. For example, it may inspect a request, retrieve information, call a business system, compare results, and ask for approval before completing an action. The defining feature is not conversational language; it is the combination of reasoning with tool use and state.

The concepts behind intelligent agents are valuable because they force teams to think about goals, observations, actions, and constraints. An agent should be selected because the workflow benefits from adaptive multi-step behavior—not because agents are currently fashionable.

Agents should also have explicit stopping and escalation rules. A workflow that cannot complete a tool call should not retry indefinitely, invent missing facts, or silently switch to a more privileged path. Clear limits on retries, approval points, and failure messages make the system easier to operate and reduce the chance that adaptive behavior becomes unpredictable behavior.

Model selection follows from constraints as much as capability

Two models may both be able to complete a task but differ in latency, cost, deployment options, data handling, context length, output quality, or safety behavior. The “best” model depends on the workload. A small, fast model may be preferable for high-volume classification; a more capable model may be justified for complex reasoning; a specialized service may outperform a general model for a narrow extraction task.

This is where current Microsoft Azure AI knowledge should be organized around capabilities rather than a memorized service catalog. Platforms evolve. The decision criteria—input type, required output, quality, risk, latency, cost, and integration—are more durable.

Cost and latency belong in the problem frame as well. A capability that works in a demonstration may be impractical at production volume if every request invokes a large model or performs several retrieval and tool steps. Estimating request frequency, acceptable response time, and value per transaction can influence whether the design uses a smaller model, a cached result, a deterministic rule, or AI at all.

The problem frame also determines the responsible-AI controls

Different workloads create different risks. A vision system can perform unevenly across lighting conditions or demographic groups. A language classifier can encode bias in labels. A generative system can fabricate claims. An agent can take an unwanted action. A responsible design process therefore asks how the specific workload can fail and which people or assets experience the consequence.

The shift from the retired AI-900 to current AI-901 reinforces this practical mindset. Beginners still need to recognize vision, language, and generative workloads, but they are increasingly expected to connect those labels to implementation. The right starting question is not “Which AI service should we use?” It is “What problem are we solving, and how will we know the result is good enough?”

Related Posts

• How Attack Paths Form Across Enterprise Systems

• Azure RBAC: Separate Scope From Role

• Azure Backup and Site Recovery Protect Against Different Failures

• Subnetting Gets Easier When You Stop Memorizing Tables

• DHCP and DNS: Two Services That Make Everything Else Look Broken

• REST APIs for Network Engineers Who Grew Up on the CLI

• Observability for AI Systems: What to Measure Beyond Latency

• Event-Driven GenAI: Where Serverless Fits

• QoS Manages Congestion, Not Speed

• Diagnosing Enterprise Routing Failures