AI & Machine Learning
Microsoft AI-103: Secrets Management for AI Apps
Secrets management for AI applications starts with a simple priority: do not create a secret when identity-based authentication can solve the same problem. Managed identities and Microsoft Entra ID can remove API keys and client secrets from many Azure-to-Azure connections. When a workload still needs a secret—for a third-party API, legacy service, or connection that cannot use Entra ID—Azure Key Vault provides controlled storage, access, and rotation. Microsoft Foundry also supports Key Vault-backed connections for scenarios that require stored connection secrets. Current documentation notes important operational limits around bring-your-own Key…
Microsoft AI-103: Responsible AI Reviews on Azure
A responsible AI review should happen while architecture can still change, not after a system is already politically or operationally difficult to stop. Microsoft frames responsible AI around six principles: fairness, reliability and safety, privacy and security, inclusiveness, transparency, and accountability. A useful review turns those principles into concrete questions about the application’s users, data, actions, failure modes, and controls. Current Microsoft guidance for agents treats responsible AI as a release gate scaled to risk. Foundry also provides risk and safety evaluations for categories such as hateful and unfair content,…
Microsoft AI-103: Reranking for Better Azure RAG
Reranking improves RAG when the retriever finds the right documents but does not place the best evidence high enough in the result set. Azure AI Search semantic ranker is an L2 reranking stage: it takes an initial candidate set produced by text or hybrid retrieval and applies deeper language understanding to move the most semantically relevant results toward the top. That distinction matters. Reranking cannot recover a document that the first-stage search never returned. It is strongest when recall is already reasonable and the problem is ordering. This makes reranking…
Microsoft AI-103: Reproducible ML Pipelines on Azure
Reproducibility means a team can explain and recreate how a machine learning artifact was produced. Azure Machine Learning pipelines help by breaking a workflow into versioned components with explicit inputs and outputs, while environments, data assets, registries, and job metadata preserve the dependencies around those components. Reproducibility is not identical output from every stochastic training run; it is the ability to reconstruct the conditions and evidence behind the result. Current Azure Machine Learning guidance uses SDK and CLI v2 as the active path for pipeline development. Components are self-contained steps…
Microsoft AI-103: REST API Patterns for Azure AI
REST is the lowest common denominator for Azure AI integration. It is useful when a language SDK does not expose a new capability yet, when a platform team needs one consistent integration surface across languages, or when an application already uses OpenAI-compatible HTTP tooling. The cost of that flexibility is that the application owns more details: authentication, headers, retries, streaming, versioning, pagination, and error interpretation. Microsoft Foundry exposes project-scoped REST endpoints for agents and an OpenAI-compatible Responses API. The project endpoint adds Foundry-specific platform capabilities, while Azure OpenAI endpoints remain…
Microsoft AI-103: Python SDK Patterns for Azure AI
Python is often the shortest path from an Azure AI prototype to maintainable application code, but SDK choice and client structure matter once the project grows. Microsoft Foundry exposes project-scoped APIs, model inference, agents, evaluations, and platform tools through several Python surfaces. The most useful pattern is to choose one abstraction deliberately instead of mixing portal-generated snippets, low-level HTTP calls, and several SDK generations in the same application. Current Microsoft documentation identifies azure-ai-projects as the Python package for Foundry project operations and supports OpenAI-compatible access through the Foundry project endpoint….
Microsoft AI-103: Protecting RAG from Poisoned Data
RAG systems can be attacked through the data they retrieve. An attacker does not always need to break the model or compromise the application directly; they may only need to place malicious, misleading, stale, or adversarial content into a source that the retriever trusts. If that content ranks highly, the model can treat it as grounding evidence or even as an instruction. This threat includes classic data poisoning, indirect prompt injection, manipulated documents, untrusted external sources, and compromised ingestion pipelines. Azure AI Content Safety Prompt Shields can detect document attacks,…
Microsoft AI-103: Prompt Versioning in Azure AI
Prompt versioning is necessary because prompt text is production behavior. A change to system instructions, examples, tool guidance, or retrieval framing can alter quality, safety, latency, cost, and tool use without any application-code change. Teams that edit prompts directly in a portal without preserving history lose the ability to explain why behavior changed or to restore the last known-good configuration. Current Microsoft Foundry provides immutable versions for prompt agents: after a saved version is created, later edits are saved as a new version. Foundry’s code-first path also allows agent definitions…
Microsoft AI-103: Prompt Injection Defenses on Azure
Prompt injection is the attempt to make an AI system follow instructions that conflict with the developer’s intended behavior. The attack can come directly from the user or indirectly from content the system reads, such as documents, websites, emails, search results, or tool output. Indirect injection is especially important for RAG and agents because untrusted content can enter the model as if it were evidence. Azure AI Content Safety Prompt Shields detects user prompt attacks and document attacks. Detection is useful, but it is one layer in a broader defense….
Microsoft AI-103: Private Networking for Azure AI
Private networking for Azure AI is about removing unnecessary public exposure without confusing network isolation with identity or authorization. A private endpoint changes how traffic reaches a service. It does not decide who is allowed to call the service, which model a workload may use, or which data an agent can retrieve. Production architecture needs both network controls and identity controls. Microsoft Foundry can be configured with public network access disabled and private endpoints for inbound access. Supporting services such as Azure AI Search, Storage, Key Vault, Cosmos DB, and…
Microsoft AI-103: Online Evaluation for AI Systems
Online evaluation closes the gap between a controlled benchmark and the behavior users actually experience. Prelaunch datasets are essential for release gates, but production traffic introduces new phrasing, unexpected tool paths, changing knowledge, rare edge cases, and interactions the development team never wrote down. A mature AI system therefore needs a way to measure deployed behavior without treating every quality issue as a support anecdote. Microsoft Foundry can evaluate deployed interactions and Application Insights traces with the same evaluation framework used for offline quality analysis. Some production-trace and conversation evaluation…
Microsoft AI-103: Monitoring Model Drift in Azure ML
Model drift monitoring answers a production question that offline accuracy cannot: has the relationship between the model, its inputs, and the environment changed enough that yesterday’s validation no longer describes today’s behavior? Azure Machine Learning model monitoring helps teams compare production data with a reference and track signals such as data drift, prediction drift, data quality, feature attribution drift, and model performance where ground truth is available. Drift is not automatically a failure. A seasonal change, new customer segment, or product launch can move the input distribution while the model…
Microsoft AI-103: Managing Agent Memory on Azure
Agent memory is useful when an AI system needs continuity beyond one request or one session. It can remember a user’s stable preferences, summarize prior interactions, or retain reusable procedures learned across conversations. That same persistence creates governance questions: what deserves to be remembered, how long should it live, who can read it, and how can it be corrected or deleted? Microsoft Foundry Agent Service currently provides managed long-term memory in preview. The service distinguishes short-term conversational context from persistent memory and supports user profile memory, chat summary memory, and…
Microsoft AI-103: MLOps and GenAIOps Together
MLOps and GenAIOps are not competing operating models. They solve overlapping parts of the AI lifecycle. MLOps grew around training, registering, deploying, and monitoring predictive models. GenAIOps extends those practices to systems where behavior also depends on foundation-model selection, prompts, retrieval, agent orchestration, safety controls, and evaluation. Microsoft’s Azure Well-Architected guidance explicitly treats GenAIOps as a specialization that complements established DevOps, DataOps, and MLOps practices. The operational stages remain familiar: prepare data, validate behavior, automate release, observe production, and maintain the system. The assets and quality checks expand. That makes…
Microsoft AI-103: Latency Tuning for Azure AI Apps
Latency in an Azure AI application is the sum of several systems, not a single model response time. A user can wait on authentication, retrieval, prompt assembly, model queueing, time to first token, token generation, tool calls, safety checks, and network hops. Tuning only the model endpoint can leave most of the delay untouched. Microsoft’s Azure OpenAI guidance separates system throughput from per-call latency and emphasizes that output token count is usually one of the strongest latency drivers. Streaming can improve perceived responsiveness even when the total generation time is…