Practice Exams:

Claude Development

Claude application development spans more than prompt writing. Production systems have to decide how model calls are structured, how tools are exposed, how errors are handled, how repository context is managed, how agent loops are orchestrated, and how retrieved enterprise knowledge is supplied safely. Those choices determine whether an impressive prototype becomes a maintainable application.

This hub organizes practical patterns around Anthropic development and the CCDV-F pathway. It focuses on boundaries: what belongs in the model request, what belongs in application code, what belongs in tools and data systems, and how each layer can fail without confusing the others.

Treat model integration as an application boundary

A Claude request should enter the system through a small number of well-owned interfaces. Centralize model selection, request metadata, timeouts, retries, and tracing so individual features do not invent incompatible behavior. The guide to SDK design shows how a client wrapper, request builder, tool registry, and state boundary can keep responsibilities clear without hiding useful platform features.

Prompts, schemas, and tool descriptions deserve version control and tests because they influence runtime behavior. Keep business authorization and durable state outside the prompt. The model can reason over those facts, but trusted application components should remain the source of truth for permissions, identity, and records that must survive beyond a conversation.

For production architecture, the broader Claude production material complements this hub with reliability, evaluation, and operating concerns. Development and operations are connected, but separating their design questions helps teams reason about each layer.

Build tools as typed, permissioned capabilities

Tool use turns a model into an actor, so every tool needs a narrow contract and an executor that validates authority independently. The article on safe Claude tools covers schema design, least privilege, idempotency, approval boundaries, and audit trails. The model requests an action; the application decides whether and how that action can occur.

The companion tool-calling patterns guide looks at orchestration choices: simple reads, staged writes, parallel independent calls, manual loops, SDK runners, and large toolsets. Selecting a pattern based on side effects and control requirements is more reliable than giving every workflow the same generic agent loop.

Tool results should be compact and purposeful. Returning an entire database object or verbose stack trace increases context and can leak information. Shape results around the next decision while preserving detailed diagnostics in internal logs.

Make failure behavior explicit

Network calls, rate limits, overload, bad inputs, timeouts, and downstream tool failures are normal operating conditions. Claude API errors separates permanent client failures from retryable conditions, explains why request IDs matter, and treats streaming interruptions as a distinct failure mode.

Retries must account for side effects. Repeating a model request is not equivalent to repeating a tool action that may already have changed external state. Idempotency keys, transaction IDs, and checkpointed workflows help the application recover without duplicating work.

Observability should connect API behavior to product outcomes. Track request latency, retry counts, model and configuration versions, tool outcomes, and application success while protecting sensitive prompt content.

Use context progressively in large codebases

Claude Code is most effective on large repositories when it does not try to load the repository wholesale. The guide to large repositories recommends a progressive workflow: map major boundaries, keep always-on project instructions concise, search for task-specific evidence, use isolated context for broad investigations, and validate with the repository’s own tests.

Project instructions such as `CLAUDE.md` should carry stable conventions that matter broadly. Path-specific rules and skills can hold narrower guidance. This keeps high-value code and task evidence from competing with a giant instruction file on every request.

Repository permissions also matter. A coding agent may have access to build scripts that publish packages or infrastructure commands that affect real environments. Treat those as tools with explicit permission boundaries, not as harmless shell commands because they live in a source tree.

Modernize incrementally instead of rewriting on faith

Legacy systems contain undocumented contracts. legacy modernization starts with characterization tests and production evidence, then creates seams where old and new implementations can coexist. Claude can accelerate explanation and refactoring, but the team still needs evidence that behavior remains correct.

Data migrations and dependency upgrades deserve their own rollback and compatibility plans. Small slices make regressions attributable; big-bang rewrites combine too many unknowns into one release. CI, canary techniques, and reconciliation metrics provide the proof that a replacement is ready.

Temporary adapters and flags should have retirement criteria. Incremental migration is safer only if the compatibility layer is eventually removed rather than becoming another permanent architecture.

Ground answers with retrieval when private knowledge matters

When users ask questions about proprietary documents or frequently changing knowledge, retrieval-augmented generation can provide the relevant evidence at request time. RAG with Claude covers corpus design, chunking, embeddings, hybrid search, reranking, provenance, access control, and evaluation.

A vector database is not a memory replacement. Retrieval should be measured independently from generation so teams can tell whether a bad answer came from missing evidence or model interpretation. Search-result content blocks and citation-aware designs can preserve source attribution through the answering layer.

Security applies before generation. If retrieval returns a document the user is not allowed to see, no prompt can repair the privacy breach. Enforce access with trusted identity data and treat retrieved instructions as untrusted unless the application explicitly designates them as policy.

Design for the platform you operate today

Claude’s platform continues to evolve across client SDKs, the Agent SDK, Managed Agents, tools, context management, and retrieval features. Avoid copying old implementation patterns simply because they appear in an earlier tutorial. Start with current official documentation, then isolate platform-specific code behind clear application boundaries so upgrades do not require rewriting the product.

The most durable architecture is not the one with the most agent features. It is the one where model reasoning, application policy, tools, data, and operations have understandable ownership. That makes new capabilities easier to adopt because the team knows which layer should change.

Use this hub as the starting point for Claude application work, then move into the focused guides as the problem demands. The objective is practical engineering: safe tools, predictable errors, manageable context, testable SDK boundaries, controlled modernization, and retrieval systems that can prove where their answers came from.

Test Claude systems at the seams

A maintainable Claude application can be tested in layers. Unit tests cover request construction, tool validation, authorization, parsing, and domain logic. Contract tests exercise the SDK boundary and representative content blocks. Integration tests verify live credentials and platform features. Evaluation sets measure model decisions such as tool choice, grounded answering, and adherence to product policy.

Separating those layers makes failures explainable. A schema-validation regression should not be diagnosed through an end-to-end model eval, and a retrieval-quality problem should not be blamed on the SDK client. Each seam should have evidence appropriate to its responsibility.

Use production traces to expand the test set. When users expose a new ambiguity or a tool repeatedly fails on one argument shape, turn that case into a regression test. The application becomes more reliable by converting operational surprises into repeatable evidence.

Keep humans responsible for irreversible decisions

Claude can summarize evidence, propose actions, draft migrations, and coordinate tools, but high-impact decisions still need an accountable owner. Define which actions can proceed automatically, which require policy checks, and which need human approval. The boundary should be based on consequence and reversibility rather than on whether the action was suggested by a model.

This is especially important when a workflow combines private retrieval with tools. A grounded answer may be correct while the proposed action exceeds the user’s authority. Authorization and approval should therefore be checked at execution time using trusted application state.

The strongest Claude systems are not the ones that remove people from every loop. They are the ones that use model reasoning where it adds leverage while keeping data ownership, permissions, business accountability, and recovery controls in explicit system components.

Architecture decisions should also account for data retention and environment boundaries. A prototype may run with developer credentials and copied sample documents, while production needs tenant isolation, secret management, audit controls, and a clear policy for what conversation or tool data is retained. Define those rules before usage scales, because retrofitting identity and retention after many integrations exist is much harder than enforcing them at the first shared boundary.

Finally, document the operational owner for each layer. Someone should own the model configuration, someone the tool executors, someone the retrieval corpus, and someone the product-level success criteria. Ownership can live in one small team, but the responsibilities should still be named. Clear ownership keeps incidents from bouncing between “AI,” application, and infrastructure teams when the actual failure is a specific contract between them.

A final design habit is to keep examples and experiments separate from the production contract. Prototypes are valuable for discovering model behavior, but once a workflow matters to users, freeze the assumptions that the application depends on: tool schemas, authorization checks, data sources, error behavior, and evaluation cases. That transition from experimentation to explicit contract is what turns a clever Claude demo into an engineering system that other people can operate and extend.

Stream responses without turning transport into business logic

Interactive Claude applications often benefit from streaming responses, especially when answers are long or tools create noticeable pauses. Treat the stream as structured message state rather than a raw text pipe. The application should accumulate a canonical response, render partial content as a projection, and distinguish completion, cancellation, and mid-stream failure.

Streaming also changes testing. Long-lived connections, partial blocks, tool requests, and interrupted responses create failure modes that do not appear in a simple non-streaming demo. Keep cancellation and retries explicit, and make sure the user interface cannot accidentally turn a partial response into a durable record that looks complete.

Test prompts against product behavior

A Claude prompt should move through a repeatable testing process before production. Build cases from real product requirements, include ambiguous and adversarial inputs, separate deterministic checks from qualitative grading, and score retrieval or tool behavior at the layer where it actually fails.

Use production traces to expand the evaluation set after launch. A user-reported failure becomes most valuable when it is converted into a durable regression case. Over time, the test suite should explain what the product has learned about its own edge cases instead of remaining a small collection of ideal examples.

Version prompts as deployable artifacts

Prompt changes deserve the same traceability as code changes. Version prompts with immutable identifiers, record the exact model and tool bundle used with them, attach evaluation results, and keep the last known-good version available for rollback.

Versioning is especially important when prompts are only one part of the behavior. Retrieval indexes, examples, tool descriptions, schemas, and policy components may change independently. Good traces identify those components separately so an incident can be reproduced without guessing which hidden configuration shifted.

Related Posts

• AI Infrastructure in Practice

• Anti-Money Laundering Operations

• AWS Architecture in Practice

• AWS Cloud Operations

• AWS Security Engineering

• Azure AI Engineering

• Azure Architecture in Practice

• Cisco Security Engineering