Practice Exams:

Anthropic CCA-F: Retrieval Design for Claude Applications

Retrieval design for Claude applications determines which external evidence reaches the model, how that evidence is ranked, how permissions are enforced, and whether users can verify the answer. Anthropic does not provide its own embedding model; its current embeddings guide points developers toward external embedding providers such as Voyage AI. Claude’s Messages API then provides document, citation, and search-result formats that make retrieved content easier to ground and attribute.

This separation is useful: the application owns ingestion, embeddings, vector or lexical search, filtering, reranking, and source lifecycle, while Claude owns reasoning over the evidence supplied to the request. Production RAG quality depends on both halves and should be evaluated as a data product, not merely a prompt.

Retrieval is therefore a core architecture layer inside Claude Production Engineering.

Start with authoritative sources

Decide which repositories, databases, websites, files, or APIs are allowed to define the answer.

RAG design is valuable when the application needs current or private facts that should not be left to model memory.

More documents are not automatically better; duplicate, obsolete, or conflicting sources can make retrieval less trustworthy.

Choose embeddings as a separate dependency

Anthropic’s current documentation explicitly states that Anthropic does not offer an embedding model and illustrates retrieval with Voyage embeddings.

Evaluate embedding providers by domain quality, dimensions, latency, price, language coverage, and migration cost.

Vector design should preserve model/version metadata so the index can be rebuilt when the embedding representation changes.

Chunk around user questions

The retrieval unit should preserve the evidence users need: a policy section, API operation, support procedure, product record, or other meaningful boundary.

Chunking evaluation should include known-relevant passages and difficult boundary cases.

Large chunks preserve context but dilute similarity; tiny chunks improve precision but can separate a statement from the exception that changes its meaning.

Use metadata for hard filters

Tenant, department, product, language, date, document status, sensitivity, and entitlement should remain structured attributes rather than being inferred from embeddings.

Filter ineligible content before it reaches Claude.

Hybrid retrieval can combine exact or lexical signals with semantic ranking when identifiers and natural-language meaning both matter.

Return search results with provenance

Claude’s current search_result content blocks let custom retrieval tools return source, title, text blocks, and optional citation settings.

When citations are enabled, Claude can attach citations to the claims that draw on those results.

This makes source identity part of the API contract instead of asking the model to manufacture footnotes from plain prompt text.

Use citations where users must verify

Claude also supports citations over supported documents and custom content.

Citation granularity depends on how the source content is chunked, so smaller focused result blocks can produce more precise attribution.

For high-consequence answers, make the source accessible enough that a user or reviewer can inspect the original evidence rather than trusting the generated explanation alone.

Rerank when first-stage search is broad

A vector or hybrid search can retrieve plausible candidates quickly, while a reranker or model-based second stage can improve ordering.

Use reranking when the correct source regularly appears but too low in the list.

Measure the quality gain against added latency and cost; do not add a second model call to every query simply because reranking is available.

Cache only stable retrieval

Repeated questions over the same approved evidence can benefit from caching, including caching search-result blocks or source documents for Claude requests.

Prompt caching should include corpus or permission version in the invalidation design.

Never reuse a cached result across users whose source eligibility differs.

Evaluate retrieval and generation independently

A bad answer can come from missing evidence, wrong ranking, stale content, or Claude misusing correct evidence.

Claude evaluation should score retrieval recall or relevance separately from answer correctness and citation quality.

Fix the layer that failed instead of immediately changing the model.

Dynamic retrieval tools should return concise, high-signal results. Anthropic’s current search-result guidance recommends returning only the most relevant results to avoid context overflow. Retrieval quality often improves when the application removes low-value candidates before they enter the context rather than asking Claude to sort through dozens of weak passages.

Search-result content can come from client tools or be provided directly as prefetched content. The choice depends on whether the application needs live search during the conversation or already has a result set from an upstream service. Keep source formatting consistent so citations and debugging remain predictable across both paths.

Retrieval failures need explicit behavior. If no eligible source supports the answer, the application can ask a clarifying question, state that it lacks evidence, or hand off to a human. Falling back silently to Claude’s general knowledge can undermine a product whose value proposition is grounded enterprise information.

Deletion and permission changes must reach the retrieval index. Removing a source file is insufficient if embeddings, cached search results, or derived summaries remain queryable. Track source IDs through the ingestion and cache layers so privacy, legal, or security removal can be completed end to end.

Observability should connect the user request to filters, top retrieved sources, scores or rank, citations, model, and final outcome. GenAI observability becomes actionable when support teams can tell whether Claude saw the wrong source or saw the right source and reasoned poorly.

A mature retrieval system is intentionally explainable: authoritative corpus, versioned embeddings, evaluated chunking, hard eligibility filters, measured ranking, visible provenance, and lifecycle. Claude can then focus on synthesis while the application remains accountable for the evidence it chose to provide.

Embedding choice should be evaluated independently from Claude model choice. A strong generation model cannot recover evidence the retriever never surfaced. Build a retrieval benchmark with query/document relevance labels and compare embedding models, chunking, metadata, and index settings before using answer quality as the only signal.

Lexical signals remain valuable for exact identifiers, codes, names, and domain-specific terms. A hybrid retrieval stack can combine keyword or BM25-style search with dense embeddings, then merge or rerank candidates. This often works better than forcing every query into semantic similarity, especially in technical support and regulated domains with exact terminology.

Query rewriting can improve recall but should be visible. Claude can generate search queries, expand acronyms, or split a complex request into subqueries. Record those transformations so operators can tell whether a miss came from the user’s wording, the rewrite, or the search engine. High-risk systems should avoid rewriting that removes restrictive terms such as tenant, jurisdiction, or time window.

Reranking can use a cross-encoder, LLM judge, or domain model to improve ordering after broad retrieval. Evaluate it separately from first-stage recall. If the correct source never appears in the candidate set, reranking cannot help; if it appears consistently but at rank eight, reranking may be an efficient improvement.

Retrieval systems should define what “current” means. Some corpora change hourly, others quarterly. Store source version or update time in metadata, exclude superseded documents where possible, and design synchronization monitoring around the business freshness requirement. A citation to an obsolete policy is still a grounded answer, but it is operationally wrong.

Access control should be enforced before ranking results become model context. The search layer should receive the authenticated user or tenant boundary and return only eligible documents. Do not retrieve broadly and ask Claude to suppress unauthorized passages after the fact. Security belongs in data selection, not in a final language instruction.

Retrieval should degrade safely. If the vector store is unavailable, the application might use lexical search, a curated fallback source, or explicitly say evidence is unavailable. The fallback should never silently broaden access or switch to stale cached results whose authorization cannot be verified.

Citations should be tested as part of UX. Users need source titles that mean something, stable links or identifiers, and enough granularity to verify the claim without opening a hundred-page document and searching manually. Good retrieval design therefore includes source metadata and presentation, not only ranking quality.

Finally, keep ingestion and query configuration versioned together. A new embedding model, metadata rule, chunking pipeline, or index can change what Claude sees even when the application prompt is unchanged. Release records should identify the retrieval configuration behind each production response so regressions can be traced and rolled back.

Retrieval testing should include adversarial and ambiguous queries. Users may misspell an identifier, use an old product name, combine two intents, or ask a question whose answer appears in several conflicting documents. These cases reveal whether query rewriting, hybrid search, metadata, and reranking work together or merely optimize clean benchmark questions.

Keep retrieval ownership close to source governance. The team that knows whether a document is authoritative, superseded, restricted, or legally retained should influence ingestion metadata and deletion. RAG quality improves when knowledge management and search engineering share one lifecycle rather than operating as separate projects.

For very large corpora, retrieval architecture should include partitioning or routing before vector search. A query about payroll policy does not need to search engineering runbooks, and a tenant-specific request should not fan out across every customer’s index. Narrowing the search domain improves relevance, latency, and security simultaneously.

Review retrieval design after model upgrades. A stronger Claude model may make better use of noisy context, but that is not a reason to accept poor retrieval. Clean evidence, explicit provenance, and good eligibility filters remain valuable because they reduce cost and make answers easier to verify.

Related Posts

• Anthropic CCA-F: Choosing the Right Claude Model

• Anthropic CCA-F: Claude Agents and Human Approval

• Anthropic CCA-F: Claude Context Windows in Practice

• Anthropic CCA-F: Cost Control for Claude Workloads

• Anthropic CCA-F: Designing Claude Applications for Production

• Anthropic CCA-F: Designing Multi-Step Claude Workflows

• Anthropic CCA-F: Evaluating Claude Responses at Scale

• Anthropic CCA-F: Guardrails for Claude Applications

• Anthropic CCA-F: Latency Tuning for Claude Applications

• Anthropic CCA-F: Reliable JSON from Claude