Practice Exams:

Microsoft AI-103: Chunking Strategies for Azure RAG

Chunking determines what a retrieval system is capable of finding. In an Azure RAG application, documents are rarely useful as one large block of text. The retrieval layer needs units that are small enough to match a focused question but large enough to preserve the context that makes the answer meaningful. That balance is why chunking deserves to be designed and evaluated rather than treated as an indexing default.

Current Azure AI Search guidance supports several approaches, from fixed text splitting to structure-aware chunking with document layout information. Integrated vectorization can combine chunking and embedding inside an indexer-driven pipeline, while custom ingestion code can apply domain-specific rules before content reaches the index. The right approach depends on the shape of the source documents and the questions users actually ask.

This is part of the same retrieval discipline covered in Azure AI Search. Search quality begins before the query arrives, because chunk boundaries decide which evidence can be retrieved independently.

Start with the structure already present in the document

A document normally contains clues about how it should be divided: headings, paragraphs, list items, table boundaries, page sections, code blocks, or records. When those boundaries carry meaning, preserve them. Splitting every document into identical character counts can cut a definition away from its explanation or a table row away from its header.

Azure AI Search can use document layout information to create semantically coherent sections from structured documents. That is useful for reports, manuals, policy documents, and other content where headings and sections already organize the subject. A chunk can then represent a complete idea rather than an arbitrary slice of text.

Structure-aware chunking does not mean every heading should become one chunk. Some sections are too large and still need to be divided. The useful rule is to preserve meaningful boundaries first, then enforce model and retrieval limits inside those boundaries.

Fixed-size chunks are simple, but simplicity has a cost

Fixed-size chunking is predictable and easy to implement. A pipeline can split content by characters or tokens, optionally with overlap, and produce a uniform corpus. This can work well for homogeneous text where document structure is weak or inconsistent.

The weakness is semantic fragmentation. A fixed boundary can divide one statement across two chunks or combine the end of one topic with the beginning of another. Larger chunks reduce fragmentation but make retrieval less precise. Smaller chunks improve targeting but can remove the context needed for the model to interpret the passage.

That tension is why chunking and evaluation belong together. There is no universally correct chunk size. The optimal size is the one that performs well on representative queries for the corpus.

Overlap should preserve continuity, not duplicate the whole corpus

Chunk overlap copies a small amount of text from one chunk into the next so that concepts crossing a boundary remain retrievable. It is useful for narrative or procedural content where a sentence or paragraph may depend on what came immediately before it.

Too much overlap creates its own problems. The index grows, embedding costs increase, and search results can return several nearly identical chunks. A generator may then receive redundant evidence instead of broader context. Duplicate passages can also distort evaluation because the retriever appears to return multiple strong results when they are copies of the same source.

Use enough overlap to protect boundary meaning, then measure whether additional overlap improves recall. If it does not, the extra duplication is operational cost without retrieval value.

Chunk size should reflect the query, not only the model limit

Embedding models impose maximum input sizes, and chunking prevents truncation. That is a hard constraint, but it should not become the design target. A chunk that is technically within the model limit can still be too broad for retrieval.

Think about the typical question. If users ask for one configuration step, a six-page chunk is likely too broad. If they ask for the rationale behind a policy section, sentence-sized chunks may be too narrow. The same source document can justify different chunking strategies for different applications.

The generator’s context window also matters. Large chunks consume more prompt space, which limits how many independent sources can be supplied. Better chunking can improve both retrieval precision and the efficiency of the final generation request.

Metadata should travel with every chunk

A chunk is more useful when it retains the document context around it. Store source identifier, title, section heading, page or location, version, date, product, department, security label, and other fields that matter to filtering or citation. Those fields should come from trustworthy ingestion logic rather than being reconstructed by the model later.

Metadata supports relevance and governance at the same time. A query can filter to the correct product version, restrict results to the user’s authorized content, or boost a recent policy. The returned result can also cite a human-readable source instead of exposing an opaque vector record.

This matters especially for enterprise RAG, where RAG data governance must survive the transformation from source document to chunked index.

Different content types may need different splitters

A single ingestion pipeline can contain PDFs, Word documents, knowledge-base articles, source code, tickets, transcripts, and tabular data. Forcing all of them through one splitter produces avoidable quality loss. Code benefits from function or class boundaries. Tickets may work well as one record per issue. Tables may need row and header preservation. Long prose may benefit from paragraph-aware or semantic splitting.

Use routing in the ingestion pipeline to choose an appropriate chunking strategy by content type. Keep the output schema consistent so downstream search does not care which splitter produced the chunk.

When multiple strategies are used, include the splitter or chunking version in metadata. That makes later evaluation and migration easier because operators can identify which representation produced a bad result.

Embedding dimensions cannot rescue a bad chunk

A larger embedding model may represent a passage more richly, but it cannot restore context that the splitter removed or separate topics that were incorrectly merged. Retrieval problems should therefore be diagnosed in order: source extraction, chunk boundaries, metadata, embedding, query construction, ranking, and generation.

The design in embeddings fits after chunking, not before it. The embedding model should be evaluated on the chunks the production system will actually store.

This order also reduces unnecessary rework. If a corpus is re-chunked after embeddings have already been generated, the entire vector set may need to be rebuilt.

Evaluate chunking with retrieval metrics

Create a stable set of questions with known relevant passages. For each chunking strategy, measure whether the expected evidence appears in the top results, how high it ranks, and whether the chunk contains enough surrounding context to support a good answer. Add hard queries that target details near section boundaries.

Do not score only the final generated answer. A strong model can hide weak retrieval on easy questions, and a weak prompt can make good retrieval look bad. Retrieval needs its own benchmark.

RAG retrieval quality depends heavily on representation. Chunking is one of the highest-leverage parts of that product because it decides what the search engine can match in the first place.

Plan for re-chunking as the corpus evolves

Chunking is not a one-time migration. New document types arrive, user queries change, and better parsers become available. Keep the original source and transformation pipeline reproducible. Version chunking logic. Build indexes so a new representation can be tested in parallel before replacing the old one.

A mature Azure AI engineering workflow can treat chunking changes like code changes: run the retrieval evaluation suite, compare the candidate index, inspect regressions, and promote only when the new representation performs better for the intended workload.

For the current AI-103 path, that mindset is more valuable than memorizing one recommended chunk size. Good Azure RAG engineering starts with document meaning, tests several representations, and keeps the chunking strategy measurable and replaceable.

Parent-child indexing is another useful design when the retrieval unit and the display unit should differ. A small chunk can be embedded and searched for precision while metadata keeps the relationship to a larger parent section or source document. The application can retrieve the focused passage, then expand to nearby context only when the answer needs it. That avoids making every searchable vector large merely to preserve context.

Keep the benchmark stable across revisions so a chunking change can be compared against the same retrieval questions and source passages.

Related Posts

• Claude Development

• Claude Enterprise Operations

• Generative AI on AWS

• Generative AI on Databricks

• Generative AI on Google Cloud

• Microsoft Platform Operations

• Network Security Platforms

• Penetration Testing in Practice

• Microsoft AI-103: Building Multi-Agent Workflows on Azure

• Microsoft AI-103: Canary Releases for AI Models