Practice Exams:

Microsoft AI-103: Vector Search Design on Azure

Vector search design in Azure AI Search involves more than adding an embedding field. The index has to match the embedding model’s dimensions, choose an algorithm and similarity metric, decide whether vectors should be compressed, retain enough original data for rescoring, and balance memory, latency, recall, and storage. Those decisions become part of the search schema and can be expensive to change after a large corpus has been indexed.

Azure AI Search supports HNSW approximate nearest-neighbor search and exhaustive KNN. HNSW builds an in-memory graph for fast approximate search, while exhaustive KNN scans the vector space exactly and does not consume the same vector-index memory quota. Azure AI Search also supports scalar and binary quantization, oversampling, rescoring, narrower vector types, and dimensional truncation for supported Matryoshka embeddings.

Vector design is therefore a core part of Azure AI engineering when RAG or semantic retrieval is part of the product.

Choose the embedding first

The vector field’s dimension count must match the embedding representation used at indexing and query time. Changing the embedding model or dimensions later can require a new field or a rebuilt index.

Keep the model, dimensions, and vector profile in versioned configuration rather than burying them in ingestion code.

The broader vector design principle is to treat embeddings as part of the index lifecycle, not a hidden preprocessing detail.

Use HNSW for fast approximate search

HNSW creates a graph that supports efficient approximate nearest-neighbor search and is a common choice for low-latency vector retrieval at scale.

Parameters such as m, efConstruction, and efSearch influence memory, indexing cost, recall, and query speed. Defaults are a reasonable starting point; tune only with a representative benchmark.

HNSW consumes vector-index memory because the graph must remain available for random-access traversal.

Use exhaustive KNN as an exact baseline

Exhaustive KNN computes exact nearest neighbors by scanning the vector space. It is slower on large indexes but useful for smaller corpora, high-recall requirements, or as a benchmark for measuring approximation loss.

Because exhaustive KNN does not rely on the HNSW graph, it does not consume vector-index memory quota in the same way.

Use exact search strategically rather than assuming approximate search is always sufficient.

Match the similarity metric

Azure AI Search supports cosine, dot product, Euclidean, and Hamming metrics in appropriate scenarios. Microsoft recommends cosine for Azure OpenAI embeddings.

The index and query must use a metric compatible with the embedding model. A mismatch can make ranking meaningless even when the vectors are correct.

Record the metric with the embedding configuration so later migrations do not guess how the existing index was scored.

Compression trades precision for capacity

Scalar and binary quantization can reduce vector memory and storage. Azure AI Search can preserve original vectors for rescoring, oversample compressed candidates, and then rerank with the original representation.

This creates a practical tradeoff: use compressed vectors for efficient candidate search, then spend more work on the strongest results to recover quality.

Hybrid search and semantic reranking are separate stages; vector compression and rescoring solve a different problem.

MRL can reduce dimensions

Azure OpenAI text-embedding-3 models support Matryoshka Representation Learning, and Azure AI Search can use dimensional truncation together with quantization.

Reducing dimensions can lower vector-index size and improve query performance, but quality must be measured. Microsoft currently documents truncation together with quantization for supported configurations.

Chunking and evaluation should remain stable while compression settings are tested.

Control vector storage

Vector fields can be searchable without being returned in search responses. In many RAG systems the raw vector never needs to be retrievable.

Azure AI Search also supports reducing redundant vector storage in some configurations, but update behavior changes when vectors are not stored. Partial document updates may require the full vector to be supplied again.

Storage optimization should therefore account for ingestion and update patterns, not only query cost.

Benchmark recall, latency, and memory

A vector configuration is only good in the context of a workload. Use known relevant documents and compare recall against an exact-search baseline where practical.

Measure query latency, vector memory, index size, ingestion time, and result quality together. A setting that improves recall slightly while exhausting memory may not be sustainable.

RAG retrieval quality should remain the product metric, with infrastructure measures explaining the cost of achieving it.

Use vectors inside a larger retrieval design

Pure vector search is not mandatory. Hybrid search combines lexical and vector candidates, filters enforce metadata and authorization constraints, and semantic ranking can refine final ordering.

The strongest Azure RAG systems treat vector search as one retrieval signal among several rather than the entire search architecture.

Keep a versioned benchmark whenever HNSW parameters, compression, dimensions, embedding models, or storage strategy changes. Without the same query-and-relevance set, teams can mistake lower storage or lower latency for a retrieval improvement.

For current Azure AI certification work, the durable sequence is: choose the embedding, define dimensions and metric, pick HNSW or exhaustive KNN, test compression and rescoring, control storage, benchmark against known relevance, and integrate vectors with filters and hybrid ranking where the workload benefits.

Vector index size should be planned with growth in mind. HNSW graphs consume memory quota per partition, and the same design that works for a small proof of concept can hit service limits after the corpus expands. Estimate document count, chunks per document, vector dimensions, data type, compression, and expected growth before choosing a long-term service size. Capacity planning is easier before the index is full than during an ingestion failure.

Quantization should be tested against the queries that matter most. Binary compression can save substantial memory, but the effect on recall varies by corpus and embedding model. Oversampling plus rescoring with original vectors can recover quality, at the cost of additional query work and preserved storage. The right configuration is an evidence-based compromise rather than the maximum compression ratio.

Index schema design should preserve metadata needed for filtering and security. Vector similarity should never be responsible for enforcing tenant, product, date, document state, or user authorization. Those constraints belong in filterable fields and application policy. A vector result that is semantically perfect but unauthorized is still a bad result.

Update patterns affect storage decisions. If vectors are not stored and the application performs partial updates, it may need to resubmit the full vector. If documents change frequently, test the operational cost of re-embedding and index updates before removing storage that the update pipeline depends on. Query optimization and ingestion reliability need to be considered together.

Use exact search periodically as a quality reference even when production uses HNSW. A sampled exhaustive query can estimate how much recall the approximate configuration is losing and whether parameter or compression changes have degraded results. This creates a measurable basis for HNSW tuning instead of relying on intuition about graph settings.

Vector search should be benchmarked after major corpus changes, not only after schema changes. New document types can alter similarity distributions, introduce duplicates, or make an old HNSW configuration less effective. Retrieval evaluation should therefore be part of ingestion change control as well as search code change control.

When several vector fields exist, define why each one is needed. Separate embeddings for title, body, image, or domain-specific representations can improve retrieval, but every field adds storage and query complexity. A field should earn its place through measurable relevance improvement rather than being added because the platform supports multiple vectors.

Production dashboards should expose vector-index memory, storage growth, query latency, and retrieval quality side by side. Infrastructure saturation can appear as a relevance problem if the service is under pressure, while relevance tuning can accidentally increase resource use. Keeping both views together makes search tuning safer.

Plan migrations as parallel indexes rather than risky in-place experiments when the change is substantial. A new embedding model, dimension count, compression strategy, or chunking representation can be built in a second index and tested against the same benchmark. Traffic can move only after quality, latency, and capacity are acceptable.

Document the final vector profile in plain language as well as JSON: embedding model, dimensions, metric, algorithm, compression, oversampling, storage choice, and benchmark date. That small record helps future teams understand why the index looks the way it does before changing a parameter that was chosen for a specific tradeoff.

Run vector benchmarks with production-like filters too. Security and metadata constraints can shrink the candidate set and change recall, so an unfiltered laboratory benchmark may overstate real-world retrieval quality.

Related Posts

• AI Infrastructure in Practice

• Claude Production Engineering

• Cloud Native Infrastructure

• Google Cloud Architecture in Practice

• Production ML on AWS

• Microsoft AI-103: Capacity Planning for Azure AI

• Microsoft AI-103: Choosing Azure AI Deployment Models

• Microsoft AI-103: Grounding Azure AI with Enterprise Data

• Microsoft AI-103: Protecting RAG from Poisoned Data

• Microsoft AI-103: Testing AI Prompts on Azure