Microsoft AI-103: Threat Modeling Azure AI Apps
Threat modeling an Azure AI application starts by accepting that the attack surface is broader than the model endpoint. User prompts, retrieved documents, memory, tools, agent identities, model outputs, logs, and external services all create trust boundaries. A system can have a secure model deployment and still be vulnerable because untrusted content becomes an instruction, an agent holds excessive permissions, or a tool executes model-generated arguments without validation.
Microsoft’s current security guidance for agents recommends mapping data flows, trust boundaries, and control points before implementation. Its Catalog of AI Attack Techniques groups risks into instruction manipulation, state and memory manipulation, trust and supply-chain abuse, execution and autonomy abuse, and resource or extraction abuse. That framing turns “AI risk” into concrete attack paths engineers can test.
Threat modeling is therefore a practical part of Azure AI engineering, not a security document produced after the design is fixed.
Map the full data flow
Start with the path from the user or event source through the application, model, retrieval layer, memory, tools, and downstream systems. Mark where data is stored, transformed, retrieved, or sent outside the immediate service boundary.
Each transition is a potential trust boundary. A retrieved document can be untrusted even inside an enterprise repository. Tool output can be malformed. Persistent memory can preserve bad state across sessions.
AI data security is easier to reason about when every input, store, and output is visible on one architecture diagram.
Separate direct and indirect prompt attacks
Direct prompt injection comes from the user; indirect injection arrives through documents, websites, emails, database records, or tool output. Both can manipulate model behavior, but their control points differ.
Prompt injection defenses should scan and label untrusted content, keep data separate from developer instructions, and assume some attacks will bypass detection.
The threat model should ask what happens after a bypass. If the model is manipulated, what data can it see, what tools can it call, and which actions still require deterministic checks?
Model memory as a sensitive store
Modern agents can persist summaries, preferences, procedural memories, cached retrieval results, and other context that influences later behavior. That state can leak confidential information or become poisoned.
Agent memory needs read and write boundaries, retention, correction, and deletion. Identify who can create memory, which user or agent can retrieve it, and whether malicious input can make the system remember unsafe instructions.
Persistent AI state deserves the same access and audit thinking as a database.
Tool access is an execution boundary
Tools turn language into real-world effects. The key questions are which tools exist, which identity calls them, what arguments they accept, and whether a user or agent can trigger irreversible work.
Tool control should narrow capability through typed schemas, allowlists, server-side validation, and approval gates for high-impact actions.
The threat model should include tool latency and availability too. A failing tool can create retry storms, inconsistent state, or unsafe fallback behavior even without an attacker.
Identity limits blast radius
Shared broad credentials make every agent failure more dangerous. Use workload or agent identities that match the role and assign only the required permissions.
Agent identity and agent permissions define the maximum impact of a compromised reasoning step.
Separate deployment authority from runtime authority so a hijacked agent cannot reconfigure its own model, network, or tool permissions.
Grounding data can be poisoned
A RAG index is part of the attack surface. An attacker may add misleading content, hide instructions in a source, exploit weak source governance, or flood the corpus with near-duplicates that rank above trusted material.
RAG poisoning controls should include provenance, ingestion scanning, authorization, trust metadata, versioning, and a clean rebuild path.
Threat modeling should ask not only whether retrieval can find relevant evidence, but whether the evidence that ranks highly deserves to influence the model.
Outputs can become new inputs
Model output is untrusted whenever another system interprets it as HTML, SQL, a command, a URL, a tool action, or executable code. Insecure output handling is how an AI attack can become a conventional application vulnerability.
Validate and escape outputs before they enter interpreters. Do not execute generated commands merely because the model returned syntactically valid text.
Zero-trust architecture applies after generation too: each downstream action still has to prove it is allowed.
Resource abuse is a security risk
Attackers can exhaust tokens, requests, search capacity, tool quotas, or expensive agent loops. Availability and cost become security properties when unrestricted consumption can create a denial-of-service or wallet attack.
Use quotas, rate limits, bounded loops, timeouts, and concurrency controls. AI rate limits should be tied to workload identity and product priority where possible.
Monitoring should reveal unusual growth in tool calls, memory writes, denied actions, or retries rather than waiting for the monthly bill.
Turn threats into release tests
A useful threat model produces engineering work: controls, owners, test cases, and residual risk. High-priority threats should appear in evaluation datasets, security tests, runbooks, and release gates.
Evaluation datasets can preserve prompt attacks, poisoned retrieval cases, unsafe tool requests, and abstention scenarios so future changes do not silently remove a defense.
Review the threat model again when the application gains a new tool, memory store, data source, model, user group, or execution permission. AI systems change quickly, and the trust diagram should change with them.
For the current AI-103 path, the durable practice is to threat-model the whole AI system: inputs, context, identity, memory, retrieval, tools, output handling, resources, and operations. The model is one component inside that trust architecture.
Threat modeling should also distinguish compromise from misuse. A legitimate user may ask an agent to do something outside policy without technically exploiting the system, while an attacker may manipulate context to make the same action happen indirectly. Both paths can end at the same tool call, but the evidence and control strategy differ. Business authorization, user intent, and security validation should therefore appear as separate checks rather than being collapsed into “the model decided.”
Supply-chain risk extends beyond packages. Models, extensions, MCP servers, connectors, external APIs, and prompt libraries can all change independently from the application. Record ownership, version, update policy, and review status for those dependencies. A trusted tool that later changes its behavior can create a rug-pull style risk even when the agent configuration is unchanged.
Multi-agent systems add trust boundaries between agents. One agent’s output should not automatically become trusted instructions for another. Use typed handoffs where possible, restrict which agents may call one another, and preserve the caller identity through orchestration. Shared memory or shared toolboxes can amplify a compromised agent if every participant has the same broad access.
Human factors belong in the threat model too. Operators may over-trust confident output, approve unsafe actions under time pressure, or paste sensitive data into a system that was not designed for it. Controls should include clear user communication, sensible defaults, approval context, and runbooks for suspicious behavior rather than assuming technical boundaries alone will prevent harm.
Finally, threat models should be maintained as the system evolves. A new browsing tool, memory feature, RAG source, model upgrade, or permission can introduce an attack path that did not exist at the previous review. Tie threat-model updates to architecture changes and significant releases. The document becomes valuable when it drives design and testing continuously, not when it merely records what the system looked like at launch.
Threat models should include data deletion and revocation paths. If a document is removed, a memory is deleted, or a user’s access changes, cached context, vector indexes, trace stores, and downstream systems may still retain copies. Identify where derived data persists and how quickly the system can stop surfacing it. Security boundaries weaken when deletion applies only to the original source.
Testing should include degraded and adversarial conditions at the same time. A model under rate pressure, a failing tool, or a stale knowledge source can push the workflow toward fallback behavior an attacker may exploit. Exercise those combinations deliberately so emergency paths are not less secure than the normal path.
Record residual risk with an owner and review trigger. Some risks may be accepted for a limited pilot or read-only workflow, but that decision should expire when the system gains more users, data, or autonomy. Threat modeling is strongest when accepted risk is scoped, visible, and revisited instead of becoming permanent through inertia.
Keep one concise threat register that links each major risk to its control, owner, evidence, and next review date.