Anthropic CCA-F: Memory Patterns for Claude Agents
Agent memory is the information a Claude application preserves beyond the immediate working context so a long-running or recurring agent can remember useful facts without replaying every prior turn. Anthropic’s current memory tool is a client-executed tool that lets Claude create, read, update, and delete files under a memory directory while the application controls the actual storage. It is designed to work with context editing and server-side compaction so active context stays focused while important information can survive summarization or session boundaries.
The central design question is not “how much can the agent remember?” It is which information deserves persistence, who owns it, how it is retrieved, how long it lives, and what should never be stored. Good memory improves continuity without turning the agent into an uncontrolled secondary database.
Memory is therefore a state-management concern inside Claude Production Engineering.
Separate working context from persistent memory
The context window holds information needed for the current reasoning step. Persistent memory holds selected facts that may matter later.
Claude context should remain curated instead of carrying every historical detail simply because it might be useful someday.
Use memory to retrieve just-in-time context when needed rather than hydrating the full user history on every request.
Store durable facts, not raw conversations
Preferences, project conventions, recurring constraints, stable identifiers, and summarized progress can be good memory candidates.
Raw transcripts, temporary tool output, long error logs, and secrets usually create more risk than value.
Before writing a memory, ask whether the same information belongs in the application’s authoritative database instead.
Use structured application state for exact values
Approval status, account balances, order IDs, authorization decisions, workflow steps, and other exact business state should live in typed systems with validation and concurrency control.
The memory tool is useful for agent-readable knowledge, not as a replacement for transactional storage.
Multi-step workflows become more reliable when deterministic state and conversational memory remain separate.
Pair memory with compaction
Anthropic’s current context-management guidance recommends compaction for long conversations approaching the context limit.
Compaction can summarize the conversation, while memory preserves facts that must survive summarization or a later session.
This combination avoids the need to keep a massive transcript active simply to preserve a few important details.
Use context editing for noisy tool history
Tool-heavy agents can accumulate large search results, file contents, directory listings, or debug output.
Context editing can clear old tool results while memory keeps the distilled fact that matters.
Latency tuning also benefits because the agent stops reprocessing stale tool output on every turn.
Implement memory storage defensively
Anthropic’s memory tool is client-side, so the application is responsible for path handling, storage, encryption, access control, quotas, and retention.
The current documentation specifically warns about path traversal and recommends ensuring operations remain under the allowed memory root.
Memory handlers should reject paths that escape the directory and cap file size or result size to prevent abuse.
Protect sensitive information
An agent may encounter credentials, personal data, customer secrets, or regulated information during work.
Claude guardrails should include validation before sensitive content is written into long-lived memory.
Apply the same classification, deletion, tenant isolation, and retention rules to memory that the organization applies to other derived data.
Retrieve memory deliberately
The memory tool can list or read information on demand, which supports just-in-time context rather than always-on injection.
Use clear memory organization so the agent can discover relevant notes without scanning an ever-growing directory.
Stale or conflicting memories should have update and deletion paths; otherwise the agent can repeatedly rehydrate outdated assumptions.
Measure whether memory improves the product
Evaluate continuity, correctness, retrieval cost, privacy, context size, and failure recovery with and without memory.
For Claude agents, the durable memory pattern is durable fact → controlled write → scoped storage → just-in-time read → context compaction → review/expiry. The goal is not maximum recall; it is the smallest trustworthy memory that makes long-running work more useful.
Memory should also have ownership. A product team needs to know which memories belong to the user, organization, agent, or one workflow. Shared organizational memory can be useful for conventions, while personal memory requires stronger isolation and user-facing lifecycle decisions.
Use explicit expiry where information naturally becomes stale. A one-week project status, temporary travel preference, or incident hypothesis should not survive forever merely because no one deleted the file. Periodic cleanup reduces both security risk and retrieval confusion.
Memory migration matters when schemas or agents change. If a new agent interprets stored notes differently, transform or version the memory rather than assuming old text remains compatible. Long-lived state is a product contract just like an API or database schema.
Finally, include memory failure in testing. The store can be unavailable, corrupted, empty, or missing an expected file. The agent should degrade gracefully and avoid inventing remembered facts. A production memory system is trustworthy when absence is handled explicitly rather than hidden by plausible generation.
Memory organization should reflect retrieval needs. A single giant notes file is easy to write and expensive to search; thousands of tiny files are hard to govern. Group durable facts by user, project, topic, or workflow with predictable names and concise summaries. The model should be able to find a relevant memory from directory metadata before opening large files.
Applications should distinguish remembered user preference from instruction authority. A memory saying “the user usually prefers CSV” can guide presentation, but a memory should not override current security policy, tool permission, or explicit user instruction. Treat memory as contextual evidence lower in the trust hierarchy than current system policy and authenticated authorization.
Memory writes should be conservative. If the model infers a preference from one ambiguous interaction, storing it permanently can make future behavior worse. Consider requiring repeated evidence, explicit confirmation, or application-side rules before durable personalization is written. Incorrect memory can be more harmful than no memory because it looks like trusted historical context.
Multi-agent systems need memory boundaries. A research agent may need project notes, while a finance agent should not automatically read personal support history. Use separate namespaces or storage roots, and pass specific memory snippets between agents when the workflow justifies it. Shared memory should be a deliberate collaboration surface, not a global folder every agent can inspect.
Deletion and correction should be first-class operations. Users or business owners may need to remove a stale preference, incorrect fact, or sensitive note. The application should know which memory file contains the information, update dependent summaries, and avoid reintroducing deleted content from an old conversation snapshot.
Memory observability should focus on reads and writes rather than private reasoning. Track which memory keys were accessed, modified, or created for an important workflow, along with the agent and session ID. This is enough to investigate many continuity problems without storing a detailed internal reasoning trace.
The safest long-running agent treats memory as a governed cache of useful facts: scoped, minimal, reviewable, expirable, and subordinate to authoritative systems. That design creates continuity while preserving the ability to correct the past when reality, permissions, or user intent changes.
Memory should not silently merge identities. If a customer account changes owner, a project is reassigned, or a device is shared, the application must know whether prior memory follows the user, the organization, the resource, or the session. Namespace design and access checks should make this answer explicit before personalization data is returned.
Testing should include conflicting memories and changed facts. Ask whether the agent prefers the newest authoritative value, surfaces uncertainty, or continues using an older note. A reliable memory layer needs conflict-resolution rules just as any replicated data system does.
For long-lived production agents, review memory storage like any other data store: backup requirements, encryption, access logs, retention, residency, deletion, incident response, and migration. The interface may look like files under /memories, but the business consequence comes from the information persisted behind that interface.
Memory retrieval should also be observable for quality. Track how often a memory is used, whether the user corrects it, and whether it improves task success. Memories that are rarely useful or frequently contradicted should be removed or moved back into on-demand authoritative lookup.
For enterprise agents, separate organization knowledge from user personalization. Policies, product facts, and project conventions can be shared under controlled ownership, while personal preferences and sensitive notes should remain isolated. Mixing these categories makes deletion and access review much harder.
A good memory design stays intentionally modest. Persist only information with clear future value, retrieve only what the current task needs, and prefer authoritative systems whenever exact current state matters. That keeps continuity useful without turning memory into an uncontrolled source of truth.
Review memory ownership, retention, and correction paths whenever the agent gains a new user population, shared workspace, or long-running capability so persisted context does not outgrow the governance model around it.
Include memory absence and stale-memory cases in regression tests so Claude can ask for current information or query an authoritative source instead of inventing continuity when stored context is missing or outdated.
Keep memory reviews recurring.