RAG on Databricks: The Hard Part Is the Data
Retrieval-augmented generation can look deceptively simple: split documents, create embeddings, retrieve a few chunks, and pass them to a language model. In a production system, the language model is often the easiest component to replace. The difficult part is building a knowledge pipeline that is complete, current, permission-aware, searchable, measurable, and maintainable. That is why the current Databricks Certified Generative AI Engineer Associate exam and its Generative AI Engineer Associate certification emphasize source-document quality, chunking, Delta tables in Unity Catalog, retrieval evaluation, reranking, vector search, deployment, governance, and monitoring.
A RAG system can fail even when the model is excellent. The answer may be wrong because the needed source was never ingested, because a parser dropped a table, because a chunk separated a heading from its explanation, because permissions exposed the wrong document, or because an index was not refreshed after the source changed. Those are data-system failures expressed through natural language.
The engineering mindset is therefore to treat retrieval as a product with its own quality objectives. Before tuning prompts, teams should prove that the right evidence exists, can be found for representative queries, and remains governed from source through response.
RAG quality begins with source selection
The fundamentals in information retrieval apply before embeddings ever appear. A retriever cannot return knowledge that is missing from the corpus. Teams should identify which documents are authoritative for each question type, how frequently they change, which versions are valid, and whether important answers live in structured systems rather than documents.
Coverage should be tested against real user questions. If shipping estimates come from a transaction system, adding more policy PDFs will not solve the gap. A strong RAG design distinguishes static knowledge, operational facts, and calculated data, then chooses the right retrieval tool for each source rather than forcing every problem into a vector index.
Parsing determines what the retriever is allowed to know
PDFs, presentations, HTML, images, and office documents contain structure that can be lost during extraction. Headers, tables, footnotes, captions, and reading order affect meaning. A parser that produces clean-looking text can still discard the exact relationships a user later asks about. Extraction should therefore be validated on difficult source types, not just on simple text documents.
The broader data-quality mindset is useful because ingestion errors are data defects. Teams should measure empty output, unusually short documents, duplicate text, failed character decoding, unsupported formats, and parser changes. A broken extraction pipeline should fail visibly instead of silently feeding low-quality chunks into the index.
Chunking is a retrieval decision, not a formatting preference
Chunk size and overlap control the unit of evidence the retriever can return. Chunks that are too small may separate a statement from its qualifier; chunks that are too large can dilute relevance and consume excessive context. Document structure matters: legal clauses, troubleshooting procedures, API references, and narrative guides may need different segmentation strategies.
The current Databricks exam outline makes this explicit by connecting chunking to model constraints and retrieval evaluation. Teams should test multiple strategies against representative queries and measure whether the retrieved chunks contain the evidence needed for a good answer. The correct chunk size is the one that improves the application, not the one that follows a universal token rule.
Delta tables create a maintainable boundary between ingestion and search
Persisting processed chunks in governed Delta tables gives the pipeline a stable data layer. The table can carry document identifiers, chunk text, source path, update time, permissions, section metadata, and any fields needed for filtering or debugging. This makes the indexing step repeatable and gives teams a place to inspect what the retriever actually sees.
The architecture ideas behind a data lake are useful here: raw source, processed representation, and serving index serve different purposes. Keeping those layers distinguishable makes reprocessing, backfills, quality checks, and migration to a new embedding model much easier than treating the vector index as the only copy of the knowledge base.
Embeddings do not remove the need for metadata
Semantic similarity is powerful, but metadata provides precision and control. Product, region, document type, effective date, customer, security label, or language can narrow retrieval before similarity is evaluated. Good metadata also helps investigations because a team can see why a chunk was eligible for a query.
Metadata should be derived consistently and governed like other data. A field that is wrong or missing can exclude the correct chunk just as effectively as a bad embedding. Retrieval debugging therefore needs visibility into both semantic scores and filter decisions.
Retrieval needs its own evaluation metrics
Judging only the final answer makes diagnosis slow. If an answer is wrong, the team should know whether the retriever failed to return relevant evidence, whether ranking placed it too low, whether the prompt ignored it, or whether the model contradicted it. Build retrieval test sets with known relevant sources so recall, precision-like measures, ranking quality, and failure categories can be inspected independently.
The concepts in core generative-AI ideas become much more useful when evaluation is decomposed this way. RAG is not one model call. It is a chain of transformations, and quality improves faster when each stage has evidence instead of forcing the team to guess from final responses.
Reranking is valuable when first-pass retrieval is broad
Vector similarity is often optimized for fast candidate generation. A reranker can then apply a more expensive relevance judgment to a smaller candidate set. This is especially useful when many chunks use similar vocabulary or when semantic proximity alone does not capture which passage best answers the question.
Reranking adds latency and cost, so it should be justified by measured quality improvement. Teams can compare retrieval with and without reranking on the same evaluation set, then choose the smallest candidate pool and ranking strategy that meet the application target.
Freshness and governance decide whether a good answer is usable
A perfectly relevant answer can still be wrong if it came from an expired policy or a document the user was not authorized to see. Source timestamps, effective dates, deletion handling, index refresh, and access controls belong in the retrieval design. Governance should propagate from the source through stored chunks and the serving path rather than being bolted onto the generated answer afterward.
The production lesson is simple: better prompts cannot repair a knowledge pipeline that is stale, incomplete, or improperly governed. RAG quality starts with data coverage and ends with monitoring whether retrieval and answers still meet expectations as sources, users, and models change. The hard part is not attaching a vector database to an LLM; it is operating trustworthy knowledge as a continuously changing product.
Document lifecycle deserves explicit design. When a source is corrected or withdrawn, its old chunks should not remain searchable indefinitely. Pipelines need stable document identifiers, update semantics, deletion handling, and index refresh behavior. Otherwise a system can answer from obsolete text even though the authoritative repository has been fixed. Freshness tests should include not just whether new documents arrive, but whether superseded knowledge actually disappears from retrieval.
Access control can also change retrieval quality. Security trimming may reduce the candidate set differently for each user, which means evaluation should include permission-aware scenarios rather than testing only with an administrator identity. A retriever that works perfectly with broad access can fail for ordinary users because relevant sources were never granted or metadata filters are inconsistent. Governance and quality are therefore intertwined: authorized evidence must still be sufficient to answer the task.
Teams should keep a failure taxonomy for RAG. Common categories include missing source, extraction error, poor chunk boundary, metadata defect, embedding mismatch, filter error, low recall, ranking error, stale index, prompt misuse, and unsupported question. Tagging evaluation and production failures this way creates actionable trends. If most misses come from parsing tables, changing the language model is unlikely to help. If relevant chunks are consistently retrieved but ignored, the generation layer deserves attention instead.
Finally, the system needs a policy for uncertainty. When retrieval returns weak or conflicting evidence, the model should not be forced to invent certainty. Depending on the application, it can cite the available sources, ask a clarifying question, say that the knowledge base does not support an answer, or escalate to a human. A trustworthy RAG product is not the one that always responds; it is the one that knows when its data is insufficient.
Retrieval pipelines also need observability around volume and shape. A sudden drop in chunk count may indicate a parser failure; a surge may indicate duplicate ingestion or a source-system change. Distribution checks on document length, language, metadata completeness, and source coverage can catch problems before users notice answer degradation. Monitoring only query latency leaves the knowledge base itself invisible, even though that data pipeline determines what the model is able to know.
Evaluation should include unanswerable questions. A production knowledge base has boundaries, and the system should learn to recognize them. Test prompts that require missing data, unsupported time periods, unauthorized sources, or facts that belong in another system. The desired behavior might be a clarification request or a transparent limitation. This prevents optimization from rewarding a model that sounds helpful by inventing answers outside the retrieval corpus.