Practice Exams:

Microsoft AI-103: Prompt Injection Defenses on Azure

Prompt injection is the attempt to make an AI system follow instructions that conflict with the developer’s intended behavior. The attack can come directly from the user or indirectly from content the system reads, such as documents, websites, emails, search results, or tool output. Indirect injection is especially important for RAG and agents because untrusted content can enter the model as if it were evidence.

Azure AI Content Safety Prompt Shields detects user prompt attacks and document attacks. Detection is useful, but it is one layer in a broader defense. An agent that can call powerful tools or retrieve sensitive data still needs identity boundaries, permission checks, data separation, tool approval, and safe failure behavior.

The goal in Azure AI engineering is defense in depth, not confidence in one classifier.

Treat external content as data, not instructions

Retrieved documents, webpages, emails, and search results are untrusted input even when the user did not write them. A malicious document can contain hidden text telling the model to ignore its system instructions or disclose information.

Structure prompts so trusted instructions and retrieved evidence are clearly separated. Do not concatenate arbitrary content into the same role or field used for developer policy.

Prompt layers should make the trust hierarchy explicit.

Use Prompt Shields for direct and indirect attacks

Prompt Shields provides detection for user prompt attacks and document attacks. The document mode is especially relevant to RAG because the attacker may never interact with the assistant directly; they only need to place adversarial instructions inside content the assistant later retrieves.

Define what happens after detection. A high-confidence attack can block processing, strip the untrusted content, reduce tool access, or route the case for review.

Detection without an application response policy is only telemetry.

Least privilege limits the damage of a successful injection

Assume some injections will bypass detection. The next control is to make the agent’s authority narrow enough that a compromised reasoning step cannot cause unlimited harm.

Agent identity should match the agent’s job. Read-only agents should not have write roles. High-impact tools should require separate identities or explicit approval.

This is why prompt security and authorization cannot be separated.

Tool schemas should constrain actions

A generic “run command” or “call any URL” tool gives an injected prompt a broad attack surface. Narrow tools with typed parameters and server-side validation are safer.

Validate tool inputs outside the model. Restrict destinations, identifiers, ranges, and permitted operations. The model can propose an action, but trusted code should decide whether the action conforms to policy.

Agent permissions are strongest when the tool surface is narrow by design.

Approval gates protect high-impact actions

Not every tool call needs a human, but irreversible or high-risk actions often should. The approval requirement should be enforced by the workflow, not written only as a prompt instruction.

An injected model output should not be able to skip the approval state. The workflow should wait until a trusted person or deterministic policy grants permission.

Multi-agent workflows can separate planning, review, and execution so authority is not concentrated in one model step.

Retrieval permissions must be enforced before generation

Prompt injection can become a data-exfiltration problem when the agent can retrieve information the current user should not see. Authorization filters should remove unauthorized content before it enters the prompt.

AI data security should preserve user and workload boundaries through the retrieval layer. The model should never be asked to “remember not to reveal” content it should not have received.

This also reduces the information available to an attacker if an injection succeeds.

Memory writes should be controlled

Persistent memory creates another injection target. A malicious interaction may try to store instructions that influence future sessions.

Agent memory should distinguish durable user facts from transient text and should validate what can be remembered. High-impact or unusual memory writes may require confirmation.

Deletion and correction must work so a poisoned memory does not remain a hidden source of behavior.

Evaluate attacks as regression cases

A defense is only useful if it survives realistic attacks. Add direct jailbreaks, hidden document instructions, malicious tool output, conflicting retrieved content, and encoded or obfuscated instructions to the evaluation suite.

Evaluation datasets should include cases where the correct behavior is refusal, safe continuation without the malicious content, or escalation.

When a new production attack is discovered, convert it into a safe regression case.

Prompt injection is a system threat

No prompt can guarantee immunity. Strong systems combine Prompt Shields, trust separation, least privilege, retrieval authorization, narrow tools, approval gates, memory controls, and traceable workflows.

Content safety provides useful detection, while zero-trust architecture limits what a compromised step can reach.

For current AI-103 work, the durable lesson is to assume untrusted instructions will eventually reach the model and to design the surrounding system so one successful injection cannot become an unrestricted business action.

Instruction hierarchy should be visible in the implementation, not only in documentation. System or developer policy belongs in trusted configuration. User content belongs in a user role or equivalent input field. Retrieved documents and tool output should be passed as data with explicit delimiters and labels. The more these sources are collapsed into one text blob, the easier it is for untrusted content to impersonate higher-priority instructions.

Outbound data controls matter because indirect prompt injection often aims to exfiltrate information. Restrict which domains and services tools can call, validate URLs server-side, and avoid tools that accept arbitrary destinations from model output. Even when the model is manipulated, the network and tool layer should make exfiltration difficult.

Secrets should never be placed in prompts simply because the model is instructed not to reveal them. Use managed identity, server-side credentials, and tool proxies so sensitive tokens remain outside model context. The safest secret is one the model never receives.

Red-team tests should include multimodal and encoded content where the application supports it. An attack can be hidden in a document, markup, code block, image-derived text, or other content transformed by the ingestion pipeline. Defenses should be tested after the same extraction process used in production.

Finally, incident response should preserve the malicious input, the detected signals, the agent version, the tool path, and the final action in a controlled evidence store. That allows the team to improve both detection and architectural safeguards instead of treating each attack as an isolated prompt trick.

Use allowlists for high-risk tool actions. If an agent can send email, change infrastructure, query a database, or call an external API, the permitted operations and destinations should be constrained in trusted code. The model should not be able to expand the allowlist through natural-language reasoning.

Rate limits and anomaly detection can reduce the impact of an attack that repeatedly probes the system. A sudden increase in denied tool calls, Prompt Shield detections, or requests for sensitive resources can be an operational signal even when no single interaction proves compromise.

Keep attack telemetry separate from ordinary user analytics. Security responders may need longer retention, restricted access, and richer context for malicious sessions than product analytics requires for routine conversations.

Finally, review the defense after every new capability. Adding memory, browsing, file upload, a new MCP server, or a write-enabled tool changes the injection threat model. Controls that were sufficient for a read-only assistant may be inadequate for an autonomous workflow.

Prompt injection defenses should also protect generated content that will be consumed by another model. In multi-agent systems, one agent’s output becomes another agent’s input and can carry malicious or accidental instructions across the boundary. Treat inter-agent messages as typed data where possible, validate important fields, and do not automatically elevate the authority of text merely because another agent produced it.

Security review should distinguish successful attack detection from successful attack containment. A detector can miss an input while least privilege and approval controls still prevent harm. Conversely, a detector can flag many attacks while an over-privileged tool leaves one bypass catastrophic. Both prevention and blast-radius reduction matter.

Content sanitization can reduce noise but should not be treated as a universal defense. Removing markup, hidden text, or unsupported file features may eliminate some attack paths, yet a plain sentence can still contain adversarial instructions. Sanitize for the format risks you understand, then keep policy and authorization controls in place regardless of the cleaned text.

Products should also explain safe refusals clearly. If an attack is detected, the user experience should not leak the exact hidden system policy or sensitive detection logic. Return enough information to continue safely without teaching an attacker how to tune the next attempt.

Related Posts

• Mastering the AI-102 Exam: Your Azure AI Engineer Associate Roadmap

• AI-900 Exam Prep: Core AI Principles and Azure Integration

• Mastering AI-102: A Complete Preparation Resource

• Understanding the Core of AI-102 and the Azure AI Engineer Role

• Microsoft AI-103: Agent Identity in Azure AI Foundry

• Microsoft AI-103: Choosing Embeddings on Azure

• Microsoft AI-103: Chunking Strategies for Azure RAG

• Microsoft AI-103: Cost Control for Azure AI Apps

• Microsoft AI-103: Hybrid Search in Azure AI Search

• Microsoft AI-103: Latency Tuning for Azure AI Apps