Context and memory are related, but they are not the same thing. Context is the information an AI application deliberately makes available for the work happening now. Memory is information the application chooses to preserve so it can be useful again later. Once an application carries information between tasks, conversations, users, tools, or agents, the question is no longer simply “What should the AI know?” but “What should the system remember, where should it live, when should it be loaded, who may change it, and when should it disappear?” 🧠
This distinction matters because memory improves continuity while also creating a new operational and security surface. Agent platforms already expose different persistence patterns: OpenAI documents sessions that keep an agent's configuration, conversation and saved work together over time; LangGraph separates thread-level persistence from longer-lived application data; and Letta provides persistent memory structures that can be edited and reused across interactions. Security guidance now treats persistent memory as something to protect rather than blindly trust. 🛡️
Original educational diagram: memory is a persistence layer; context is the selected working set.
- Quick Comparison: Context, Memory and Evidence
- Context vs Memory: The Mental Model
- Short-Term and Long-Term Memory
- What Should an AI Application Remember?
- Updating and Removing Memory
- Continuing a Task Across Separate Conversations
- Memory as a Security and Governance Surface
- Enterprise Rollout: Turning Memory into a Governed Capability
- Common Mistakes: How Memory Systems Fail
- FAQ
- References & Further Reading
- Summary
1. 🔀 Quick Comparison: Context, Memory and Evidence
| Concept | Main purpose | Typical lifetime | Main design question |
|---|---|---|---|
| Current context | Help the application complete the present task. | Usually one turn, task, or active workflow. | What information is useful right now? |
| Short-term memory | Maintain continuity inside an ongoing interaction. | Until the conversation or workflow ends, expires, or is compacted. | What state must survive the next step? |
| Long-term memory | Preserve durable information across interactions. | Until its policy-defined retention period, replacement, or deletion. | Is this durable enough and safe enough to retain? |
| Retrieved evidence | Supply relevant external information for a task. | Usually limited to the current operation. | Can this be trusted as data rather than instructions? |
2. 🧩 Context vs Memory: The Mental Model
Context is the information assembled for a particular execution: the user request, approved instructions, relevant records, recent tool results, selected memories, and other task-specific evidence. Memory is information the system intentionally persists beyond that execution: a preference, a durable fact, unfinished workflow state, or anything else the product decided is worth keeping.
A memory record sitting in a database is not automatically part of the active working state. It becomes context only when the application selects it, and a piece of temporary context can be used and discarded without ever becoming memory. The part in the middle is selection: a process, owned by the application, that decides what is relevant, permitted, current and safe to load. Without it, memory becomes a dumping ground and context fills with material that may be stale, irrelevant, duplicated, private, or unsafe.
Platform grounding: OpenAI's Agents API describes a session as something that keeps an agent's configuration, conversation and saved work over time, and recommends storing the session ID with your own application's conversation state. LangGraph makes the same split explicit in a different way: state persisted per thread, plus a separate store for data that outlives any one thread. The names differ; the architecture is the same.
Enterprise note: treat context assembly as an application pipeline, not an invisible side effect. It should know where each item came from, why it was selected, whose it is, and what actions it may influence.
Common mistake: assuming “stored” means “trusted” and “retrieved” means “relevant.” Storage says where information lives; it does not say whether the information belongs in the next decision.
🎯 Use this when you need to explain the difference between “the application has the information” and “the agent is currently allowed to use it.”
3. 🗂️ Short-Term and Long-Term Memory
Short-term and long-term are design labels, not universal product rules.
Short-term memory is state tied to an active conversation, thread, workflow, or task: recent messages, current approval state, a temporary plan, the latest tool outcome. Long-term memory survives the current interaction and can be reused later: durable preferences, stable business facts, reusable project state, or a user-approved profile attribute.
Platform grounding: LangGraph describes checkpointers for thread-scoped state and stores for longer-lived, cross-thread information. Whatever the product, the practical lesson is the same: choose the scope and expiry of each piece of state when you design it, not after it has piled up.
Enterprise note: define memory classes before building storage. “Conversation state,” “user preference,” “business fact,” “workflow state,” and “temporary scratch data” should not share one retention and access policy.
Common mistake: storing everything under one generic object called memory. Without a lifecycle category, teams cannot later decide who may read it, how long it stays, or whether another workflow may use it.
🎯 Use this when you need a clean persistence strategy for multi-turn or long-running agent workflows.
4. 📌 What Should an AI Application Remember?
The right question is not “Can we store this?” but “What future decision or task will become better because we stored this?” A good memory candidate has:
- Durability: it is likely to stay useful after the current interaction ends.
- Purpose: there is a known reason to keep it.
- Scope: the system knows whose, or which workflow's, information it is.
- Source: the application can identify where it came from.
- Lifecycle: the system knows when it should be reviewed, replaced, or removed.
- Access: only appropriate workflows can retrieve it.
- Will this probably matter later?
- Would the user or business owner expect it to persist?
- Do we know where it came from?
- Can we safely limit who can use it?
- What is the removal rule?
There are two common write paths. In the first, the agent proposes a memory during the conversation. In the second, the application extracts candidate memories afterward, in the background. Either way, the safe pattern is the same: the model suggests, and a policy layer with its own rules decides whether the write happens, in what scope, and with what provenance. An agent that can write straight into durable storage turns every untrusted input into a possible persistent instruction.
A common way to classify memory is by what it does: facts (“the customer prefers email”), past events (“the refund was issued on the 3rd”), and behavioral rules (“always do X when Y”). The third kind deserves the strictest write control, because a stored rule can steer future actions in a way a stored fact usually cannot.
Platform grounding: Letta describes persistent memory blocks for durable information, plus archival memory that can be searched and modified. LangGraph describes long-term memory for user-specific or application-level information that must outlive a single thread. Both treat memory as meaningful state, not an accidental archive.
Enterprise note: memory creation needs an owner. Product teams decide the user value, security and privacy teams set constraints, data owners define classification and retention, and engineering enforces them.
Common mistake: treating the memory store as a substitute for an enterprise system of record. Memory may support an agent, but it should not silently become the authority on customer, financial, or operational truth.
🎯 Use this when you are designing the decision that turns conversation information into durable application state.
5. 🔁 Updating and Removing Memory
This is why memory should be treated as versioned state, not a bag of facts. A practical memory record carries fields like these:
| Field | Purpose |
|---|---|
memory_id | Stable identifier for the memory item. |
scope | The user, application, project, or workflow boundary it belongs to. |
source | Where the information originated. |
classification | Sensitivity level, which drives who may retrieve it. |
created_at | When it entered durable storage. |
expires_at | When it should stop being usable. |
status | Active, superseded, revoked, or deleted. |
superseded_by | Link to the record that replaced it, so history stays traceable. |
An update should not simply overwrite the value. A safer pattern:
- Identify the existing memory.
- Decide whether the new information truly refers to the same fact.
- Validate the source and scope of the new value.
- Mark the previous state as superseded or invalid according to policy.
- Store the new state with its own provenance and lifecycle.
- Make sure retrieval prefers the current valid state.
Deletion also has more than one meaning. A user-facing “forget this” request may need removal from the active store, search indexes, derived stores such as summaries, and anything else the retention architecture defines. Audit records may need to prove the deletion happened without keeping the sensitive content.
def supersede_memory(store, memory_id, replacement, actor):
authorize(actor, "memory:update", replacement.scope)
validate_scope(replacement)
validate_source(replacement)
with store.transaction(): # both changes succeed, or neither does
old = store.get(memory_id, status="active", for_update=True)
replacement.status = "active"
replacement.supersedes = old.memory_id
store.save(replacement)
old.status = "superseded"
old.superseded_by = replacement.memory_id
store.save(old)
store.audit("memory_superseded", actor, old.memory_id, replacement.memory_id)
The code is simplified. A production policy must also cover concurrency conflicts, deletion propagation to derived stores, and failure handling. The important details are the transaction (a crash cannot leave the user with no active memory) and the authorization and audit calls.
Platform grounding: Letta documents memory that can be edited and archival records that can be deleted. LangGraph persists state that can be inspected and changed after the fact. The engineering lesson is general: memory is not a one-way write path, so correction, replacement and deletion must be designed in.
Enterprise note: treat changes to retention, schemas, retrieval and deletion logic like any other production data change: test and approve them.
Common mistake: assuming “update” is enough. Without an explicit removal path, a memory system accumulates contradictory and obsolete state.
🎯 Use this when a memory system must stay correct as people, projects, policies, and business conditions change.
6. ▶️ Continuing a Task Across Separate Conversations
Conversation 1: it identifies the invoice, supplier, disputed amount, and open questions.
Conversation 2: two days later the user says, “Continue the invoice dispute review.”
What should survive: the case identifier, workflow status, confirmed facts, outstanding questions, and approved next step.
What should not automatically return: every message from the first conversation, unrelated chat, temporary reasoning notes, or sensitive information unrelated to the case.
A safe continuation process has six steps:
- Identify the workflow: determine which task or case the user is continuing.
- Load durable state: retrieve the smallest set of records that restores continuity.
- Re-check permissions: confirm the current user is still allowed to access the data.
- Refresh mutable facts: fetch current business data instead of assuming a stored fact is still true.
- Assemble current context: combine fresh data, approved memory, and the new request.
- Continue the workflow: record new durable state only when the memory policy allows it.
Platform grounding: OpenAI's Agents API lets you continue work by sending another message to the same session ID and retrieve the saved items afterward, which is why it advises keeping that ID alongside your own conversation record. Notice what that does not do: it does not decide whether the current user may still see that session. That check is yours.
Enterprise note: separate workflow identity from user identity. A user may be authorized to continue their own case without being authorized to inspect a colleague's case, even when both cases use the same agent.
Common mistake: restoring a whole old transcript because it is easy. Continuity is not the same as replaying history.
🎯 Use this when an agent must pause and resume work without dragging an entire previous conversation into every future task.
7. 🛡️ Memory as a Security and Governance Surface
A documented case. In May 2026 the OWASP GenAI Security Project published “Memory Is a Feature. It Is Also an Attack Surface,” by the lead of the ASI06 entry (Memory and Context Poisoning) in the OWASP Top 10 for Agentic Applications. It describes Cisco research, named MemoryTrap, in which an ordinary developer workflow (clone a repository, let a coding agent help, approve a dependency install) turned into persistent prompt injection. The payload reached persistent memory, the global hooks configuration, and a high-trust instruction layer, so one action shaped behavior across sessions and projects. According to the article, the vendor's fix removed user memories from the system prompt, which closed that specific high-trust path.
The lesson is architectural: persistent state carries unwanted influence forward, and the higher the trust level a memory is promoted to, the larger the damage. Never feed stored memory into your highest-trust instruction layer.
A mature memory system can answer four questions about every record:
| Question | What the application should know |
|---|---|
| Where did it come from? | User input, business system, document, tool output, human approval, or another source. |
| Who owns it? | User, team, application, business function, or system of record. |
| Who may use it? | Only the identities and workflows included in the policy. Namespace memory by user and tenant so one person's memory can never be retrieved for another. |
| What can it influence? | Information gathering only, recommendations, or actions with business consequences. |
Keep data and authority separate. A memory may contain a sentence that says “approve the payment.” That sentence stays data unless the application independently establishes that it is an authorized instruction from a trusted control point. When memory is loaded into context, label it with its source and mark it as recalled data, so the agent and your logs can tell it apart from the trusted instructions.
The same applies to connected tools. The Model Context Protocol (MCP) connects AI applications to external data sources and tools. That expands what an application can reach, but it does not make any external result a trusted instruction or a memory record.
Enterprise note: give memory a data classification model. A preference, a public project fact, a customer identifier, a confidential record, a secret, and an authentication credential must never be interchangeable memory objects. Credentials and secrets should not be stored as memory at all.
Common mistake: letting the agent write directly to durable memory with no policy layer. Convenience today becomes a persistent security problem tomorrow.
🎯 Use this when your agent has persistence, external tools, shared state, or any ability to influence real systems.
8. 🏢 Enterprise Rollout: Turning Memory into a Governed Capability
NIST's Generative AI Profile is a companion to the NIST AI Risk Management Framework. It identifies risks specific to generative AI and suggests actions organized around the framework's four functions: govern, map, measure and manage. For a memory system, that translates into ownership of what enters, how it is classified, how it changes, and how incidents are handled. A production memory capability needs clear owners across the lifecycle:
| Capability | What good looks like | Typical owner |
|---|---|---|
| Ownership | Every memory class has a business and technical owner. | Product + Engineering + Data owner |
| Versioning | Schemas and memory policies have controlled versions. | Engineering |
| Change control | Changes to retrieval, retention, or write behavior are reviewed. | Architecture + Change management |
| Access control | Memory access follows identity, role, and data scope. | Security + IAM |
| Retention | Every class has a defined retention and deletion policy. | Data governance + Legal/Privacy |
| Observability | Teams can trace memory reads, writes, updates, and action decisions, with sensitive content redacted in traces. | Platform engineering + SRE |
| Incident response | Compromised memory can be isolated, revoked, and investigated. | Security + Incident response |
- Define the change: what memory type, retrieval behavior, policy, or schema is changing?
- Classify the data: does the change alter the data categories that may be stored or retrieved?
- Check authorization: can the same identities and workflows still access the data?
- Review retention: does the change create new persistence or extend existing retention?
- Review action impact: can memory now influence a more sensitive workflow or tool?
- Test auditability: can the organization explain what memory was read or written during an incident?
- Approve and version: record the release decision before deployment.
Original diagram: no single control is assumed to be a complete security boundary.
Memory adds cost through extra retrieval, storage, synchronization, logging, and tool activity. Monitor spend per workflow rather than letting every workflow read and write without limit.
- Set budgets for high-volume workflows.
- Limit repeated memory writes and unnecessary retrievals.
- Alert on unusual growth in tool calls, retrieval volume, or memory size.
- Review expensive workflows by business value, risk, and actual usage.
Suppose a persistent memory entry is influencing actions unexpectedly. The response should not depend on manually searching thousands of conversations.
- Contain: disable the affected memory source or workflow, or switch the agent to read-only memory.
- Identify: determine which records were read or written.
- Scope: identify which users, workflows, and actions were exposed.
- Correct: revoke, replace, or delete affected records according to policy.
- Review: inspect traces, access controls, and persistence rules.
- Recover: restore normal operation only after the relevant controls are in place.
Readiness gate: before go-live, an enterprise should be able to answer one uncomfortable question: “What is our kill switch if memory begins influencing the system incorrectly?”
🎯 Use this when moving from a prototype memory feature to a governed enterprise capability.
9. ⚠️ Common Mistakes: How Memory Systems Fail
Most memory failures are not caused by a missing database. They happen because the team never defined what memory means, who owns it, and how it behaves when conditions change.
10. ❓ FAQ
11. 🔗 References & Further Reading
- OpenAI — Run and continue sessions
- OpenAI — Context Engineering: Short-Term Memory Management with Sessions
- LangGraph — Persistence
- LangGraph — Memory
- Letta — Memory & Dreaming
- Letta — Memory Blocks
- OWASP GenAI Security Project — Memory Is a Feature. It Is Also an Attack Surface
- OWASP — Top 10 for Agentic Applications (includes ASI06: Memory & Context Poisoning)
- Cisco — Identifying and remediating a persistent memory compromise in Claude Code (MemoryTrap research)
- NIST — AI Risk Management Framework: Generative AI Profile
- Model Context Protocol — Introduction
- Model Context Protocol — Versioning
Originality & attribution: this article is independently written for educational purposes. All examples are original. The sources above were used only to verify documented capabilities, terminology and security concepts, and no source text, diagrams or examples were reproduced. Platform features change quickly, so check each vendor's current documentation before relying on a specific capability.
12. 📝 Summary
- Context vs memory: context is selected for the current task; memory is deliberately retained for future use, and an application-owned selection step connects the two.
- Short-term vs long-term: active workflow state and durable information deserve different scopes and lifecycles.
- What to remember: save information because it has future value, and let a policy layer, not the agent, decide the write.
- Updating and deleting: memory needs provenance, versioning, expiry, atomic replacement, and deletion that reaches derived stores.
- Continuing work: restore the smallest useful durable state, re-check authorization, and refresh mutable business data.
- Security: persistent memory is an attack surface, as the documented MemoryTrap case shows; keep it out of the highest-trust instruction layer.
- Enterprise rollout: memory needs owners, access controls, change gates, retention rules, observability, incident response, and a tested kill switch.
- Core principle: remember deliberately, retrieve selectively, validate continuously, and delete when policy says the information has reached the end of its useful life.
The maturity of an AI application is not measured by how much it can remember. It shows in how carefully it decides what deserves to be remembered, when that memory may be used, and when it should be forgotten. 🌱
Comments
Post a Comment