Practice Exams:

Microsoft AI-103: Choosing Embeddings on Azure

Embedding choice affects retrieval quality, index size, query latency, re-indexing cost, and the long-term shape of a RAG system. It is easy to treat embeddings as an implementation detail because they sit behind vector search, but the vector dimensions are part of the index schema and the model determines how text is represented. Changing either later can require rebuilding the vector corpus.

Azure OpenAI currently supports third-generation embedding models such as text-embedding-3-small and text-embedding-3-large, alongside older options in some services. Azure AI Search can use these models through integrated vectorization or an Azure OpenAI vectorizer. The larger third-generation model supports up to 3,072 dimensions; the smaller supports up to 1,536, and both allow reduced dimensions.

The decision belongs in the same architecture conversation as Azure AI Search. A vector is useful only if it improves retrieval on the queries and documents the application actually has.

Start with retrieval tasks, not model size

Define what the embedding needs to separate. A support knowledge base may contain short technical procedures and exact product terminology. A legal repository may need semantic similarity across long formal passages. A multilingual knowledge base may depend heavily on cross-language representation. Code, image, and multimodal retrieval create different requirements again.

Create a query set with known relevant documents or passages. Include synonyms, abbreviations, misspellings, product identifiers, and ambiguous terms. The best embedding model is the one that helps the retriever place the right evidence in the candidate set consistently.

This is why embedding models must be tested on real retrieval tasks. The model is not evaluated by the beauty of the vector; it is evaluated by the quality of the search results it enables.

Large vectors are not automatically better vectors

text-embedding-3-large can emit a higher-dimensional vector than text-embedding-3-small, but more dimensions increase storage, memory, transfer size, and indexing work. Higher dimensionality may improve representation for some tasks, yet the gain has to be measured against operational cost.

The third-generation models support a dimensions parameter that lets teams reduce vector size. That creates a useful tuning axis: compare retrieval quality at several dimensions rather than assuming the maximum is necessary.

If a smaller vector preserves the required recall and ranking quality, it can reduce index size and query cost. If quality drops on the difficult cases that matter, the larger representation may be justified.

The index schema must match the embedding dimensions

Azure AI Search vector fields are configured with a fixed dimensionality. The configured field dimensions must match the vectors written into that field. For Azure OpenAI vectorizers, the maximum dimensions depend on the selected model, with the small and large third-generation models supporting different limits.

This makes embedding changes a schema-management issue. Switching from one dimension to another may require a new field or re-indexing. Teams should avoid burying model and dimensions inside application code without recording them in index configuration and deployment metadata.

Version the embedding configuration alongside the index. If the embedding model changes, keep the old index available until the new vectors have been built and retrieval evaluation passes.

Use the same representation logic at indexing and query time

Document vectors and query vectors have to live in the same embedding space. Using one model for indexing and another for queries produces meaningless similarity. Integrated vectorization in Azure AI Search can simplify consistency because the indexer and query vectorizer can reference a known deployment.

If vectorization is implemented in custom code, make the model deployment and dimensions explicit configuration. Add tests that verify vector length and model version. A silent configuration mismatch is easier to prevent than to diagnose from poor search results later.

Managed identity is useful where supported because it lets Azure AI Search call the embedding resource without storing a long-lived key. The search service identity still needs the correct OpenAI role assignment.

Cosine similarity is only one part of relevance

Azure OpenAI embeddings are commonly used with cosine similarity in Azure AI Search, but vector similarity should not be confused with final relevance. The nearest vector may be semantically related while missing an exact identifier, recent date, or important business filter.

That is why hybrid retrieval is so powerful. Hybrid search combines vector similarity with lexical matching and can add semantic reranking. The embedding model then contributes one strong signal instead of carrying the entire relevance burden.

If retrieval quality is weak, first identify whether the correct document is absent because of the embedding, chunking, filter, query, or ranking stage. Replacing the embedding model is not the universal fix.

Chunking and embeddings must be evaluated together

An embedding represents the text it receives. If a chunk combines unrelated sections, the vector becomes an average of topics the user may query separately. If chunks are too small, they may lose the context needed to distinguish similar concepts. The same model can therefore perform very differently under different chunking strategies.

Test chunk size, overlap, heading preservation, and metadata with the candidate embedding. Use real documents rather than synthetic sentences. Measure whether the returned chunk contains enough evidence for generation, not merely whether its title seems relevant.

Chunking and evaluation are often higher-leverage improvements than moving directly to a larger embedding model.

Multilingual retrieval deserves its own test set

Third-generation embedding models improved multilingual retrieval compared with older generations, but a multilingual application should still test its own language pairs and terminology. Product names, transliterated terms, abbreviations, and domain-specific vocabulary can behave differently from general benchmark datasets.

Include cross-language queries if the application expects them. A user may search in one language for a document written in another. Measure whether relevant results remain competitive and whether keyword fields can supplement vectors when exact terms matter.

Language testing should also include normalization and tokenization decisions in the ingestion pipeline. Embedding quality cannot compensate for corrupted or badly extracted source text.

Cost includes re-embedding and index lifecycle

Embedding cost is not only the price of one API call. Large corpora may need to be re-embedded when the model, dimensions, chunking, or content changes. Frequent document updates create continuous indexing demand. Larger vectors consume more storage and can affect memory and query performance.

Estimate the full lifecycle: initial corpus embedding, incremental updates, query embeddings, evaluation jobs, disaster recovery, and future re-indexing. If the corpus is very large, a modest per-document difference can become operationally significant.

That is another reason to choose the smallest configuration that meets measured retrieval requirements rather than maximizing dimensions by default.

Plan migrations before the first production index

Embedding models evolve. A production design should assume that the embedding will eventually change. Keep source documents and chunking inputs reproducible. Record which model and dimensions produced each vector field. Build the index through an automated pipeline rather than a one-time manual import.

A safe migration can create a second index or vector field, re-embed the corpus, run the evaluation set against both versions, then shift retrieval traffic after quality and latency are acceptable. This resembles blue-green releases applied to the retrieval layer.

For teams working through Azure AI developer certification, the durable rule is simple: choose embeddings with evidence. Compare candidates on your retrieval tasks, test dimensions instead of assuming the maximum, preserve hybrid search for exact language, and design the index so that the embedding model can change without turning migration into an emergency.

Recall should be established before vector cost is optimized.

Vector storage is easy to measure, so teams sometimes optimize dimensions before they know whether retrieval is good enough. Reverse that order. Establish a baseline recall target on representative queries, then test smaller dimensions or a smaller embedding model and measure the change. If the relevant passage still appears reliably in the candidate set, the smaller configuration may be a genuine efficiency gain. If recall drops on important edge cases, the storage saving is not free.

Keep the benchmark stable as the corpus evolves. Add difficult production queries to the set, but do not remove older cases simply because the new model performs poorly on them. A durable evaluation suite protects against accidental quality loss during embedding migrations and makes the tradeoff between vector size and retrieval performance visible to both platform and product teams.

Document the final embedding decision in the same place as the search schema. Record the model, dimensions, similarity metric, chunking version, evaluation date, and the benchmark result that justified the choice. Future migrations become much easier when the team can reproduce the original baseline instead of reverse-engineering an old index.

Related Posts

• Microsoft AI-103: Agent Identity in Azure AI Foundry

• Microsoft AI-103: Azure AI Content Safety in Practice

• Microsoft AI-103: Azure AI Foundry Model Selection

• Microsoft AI-103: Azure AI Search for RAG

• Microsoft AI-103: Blue-Green Releases for AI Endpoints

• Microsoft AI-103: Building Multi-Agent Workflows on Azure

• Microsoft AI-103: Canary Releases for AI Models

• Microsoft AI-103: Capacity Planning for Azure AI

• Microsoft AI-103: Choosing Azure AI Deployment Models

• Is Microsoft PL-300 Worth Getting? Everything You Need to Know