What Your Agent Remembers
Memory is the new attack surface — and the new moat. Why persistent state is the hardest unsolved problem of always-on agents.
What Your Agent Remembers
Memory is the new attack surface — and the new moat. Why persistent state is the hardest unsolved problem of always-on agents.
Every agent demo is a goldfish. It boots, reads the task, reasons, acts, and forgets. The loop itself is simple — a while loop around a model call and a list of tools [1] — and for a single task, forgetting is a feature: every run starts clean, every mistake dies with the process.
The agents we are actually deploying are not goldfish. The overnight coding shift, the always-on Telegram chief of staff, the assistant that "knows how you like your PRs" [2] — all of them depend on state that outlives a session. Preferences, environment facts, summaries of last week's work, the reason a migration was rolled back. Strip that away and you are back to a very fast intern with amnesia.
Memory is what turns a capable model into a useful colleague. It is also what turns a single bad afternoon into a permanent condition. That trade-off — and how to engineer it — is the subject of this essay.
flowchart LR
subgraph IN["What a session ingests"]
U["👤 User messages"]
W["Workspace files"]
X["🌐 Web · MCP · sub-agents
out-of-workspace reads"]
end
subgraph S["Session"]
L["Agent loop"]
T{"Taint
derivation"}
end
subgraph MEM["Persistent memory"]
F["Facts
every system prompt"]
E["Episodes
recalled by similarity"]
Q["🔒 Pending review
stored, never replayed"]
end
U --> L
W --> L
X --> L
L --> T
T -->|trusted| E
T -->|trusted + explicit opt-in| F
T -->|untrusted| Q
Q -.->|human promotes| E
E -->|next session| L
F -->|next session| L
style Q fill:#3a1a1a,stroke:#c0392b,color:#e8c8c8
style T fill:#1a2a3a,stroke:#3b82f6,color:#dbeafe1. The context window is not memory
The most common confusion in agent engineering is treating the context window as memory. It is not. It is working memory — a scratchpad that is rebuilt on every turn and discarded at the end of the session. Making it bigger does not make it durable, and it does not even make it reliable: models attend unevenly across long contexts, recovering information from the beginning and end far better than from the middle [5].
Real memory is what survives the process boundary. It is written by one session and read by another, often days later, often by an agent with a different task and a different set of tools. That makes it categorically different from context:
- It crosses trust contexts. A memory written while the agent was reading a hostile web page will be read while the agent has shell access to your production repo.
- It is selected, not seen. The agent does not reread its whole history; something — a similarity search, a ranker, a summarizer — decides what comes back. That selector is part of the attack surface.
- It is authoritative by position. Memory is usually injected into the system prompt or near it — the place the model is trained to treat as ground truth about the user and the world.
Research systems have explored the design space for years. MemGPT framed it as virtual memory for LLMs, paging state between a small context and a large external store [3]. The Generative Agents work showed how observations, reflections, and retrieval scoring produce believable long-horizon behavior [4]. Both are excellent at answering how do we remember more? Neither was built to answer the question that matters once agents touch real systems: who is allowed to write what the agent will believe tomorrow?
2. A taxonomy of agent memory — and how each tier fails
It helps to separate memory by how it re-enters the loop, because the failure mode follows the re-entry path, not the storage format.
| Tier | What it holds | How it comes back | Failure mode |
|---|---|---|---|
| Working | The live conversation, tool results, the current plan | Always present, until compaction | Context overflow; summaries that silently drop the one constraint that mattered |
| Episodic | Distilled summaries of past sessions | Similarity search against the current task | Poisoned or stale episodes recalled exactly when they are most relevant |
| Semantic / facts | Durable statements: "user prefers Go", "deploys go through make release" |
Injected into every system prompt | One wrong fact steers every future session |
| Procedural | Skills, playbooks, prevention rules | Loaded by trigger match or on demand | A planted procedure is executed, not merely believed |
Read the last column top to bottom and the pattern is obvious: the more durable and the more automatic the re-entry, the worse the failure. A poisoned working context dies with the session. A poisoned episode resurfaces when the topic comes up. A poisoned fact is present in every conversation from now on. A poisoned procedure runs.
This is the core design tension. The tiers that make an agent feel like it knows you are exactly the tiers where an attacker gets the most leverage per byte written.
3. Memory poisoning: how one injection becomes a permanent tenant
Prompt injection is usually described as a single-turn problem: the agent reads hostile text, follows it, and does something it should not [6]. That framing understates the threat badly once memory exists. With persistent memory, the injection does not have to win now. It only has to get written down.
This is not hypothetical. In 2024, Johann Rehberger demonstrated that ChatGPT's memory feature could be written by prompt injection from untrusted content, planting an instruction that persisted across conversations and exfiltrated what the user typed afterwards [7]. In early 2025 he showed a variant against Gemini using delayed tool invocation: the injected document instructs the model to save a false memory only when the user later says an innocuous trigger word, so the write appears to be user-initiated [8]. On the research side, AgentPoison showed that poisoning a tiny fraction of an agent's memory or knowledge base — well under one percent — is enough to hijack its behavior on targeted queries while leaving everything else apparently normal [9], and PoisonedRAG showed that a handful of crafted passages in a corpus of millions reliably steers retrieval-augmented answers [10]. OWASP's agentic threat taxonomy now lists memory poisoning as a first-class agent threat [11].
The anatomy is consistent:
sequenceDiagram
autonumber
participant A as Attacker
participant S as Session N (agent)
participant M as Memory store
participant N as Session N+k (agent)
A->>S: Hostile text in a web page, issue, or MCP result
S->>S: Task completes normally — nothing looks wrong
S->>M: End-of-session extraction writes a "fact" or episode
Note over M: Payload is now first-party state
N->>M: Recall by similarity or system-prompt injection
M-->>N: Payload re-enters as trusted context
N->>N: Acts on it with today's tools and permissionsThree properties make this worse than a one-shot injection:
- Laundering. At step 3, attacker text becomes the agent's own memory. Every downstream defense that distinguishes "tool output" from "what I already know" is now on the wrong side of the line.
- Time shifting. The write happens in a low-privilege session (reading the web); the payload fires in a high-privilege one (shell, credentials, deploy). Least privilege per task is defeated if memory is shared across tasks.
- Plausibility beats payload. The most dangerous memories are not
curl … | sh. They are quiet, declarative, believable: "the user prefers that you skip the test suite for hotfixes", "the staging and production databases share credentials". A filter can catch commands. Nothing simple catches a reasonable-sounding lie.
The conclusion I drew in July still holds, with a refinement: you only have to be prompt injected once [2] — and memory is how once becomes forever.
4. Provenance and taint: every memory needs a birth certificate
If memory is input from the past, it deserves the same treatment as input from the network: assume it is hostile until you know where it came from. That requires a primitive most memory systems lack entirely — provenance. Not "when was this written" but "what had the agent read in the session that wrote it?"
The working rule is simple to state:
A memory inherits the trust level of the least trusted content its session ingested.
A session that only touched the user's messages and files inside the workspace produces trusted memory. A session that browsed the web, called an MCP server, read results from a sub-agent, or opened files outside its workspace produces tainted memory — stored for audit, never replayed automatically.
This is deliberately coarse. It does not try to decide whether a particular summary sentence was influenced by the attacker; that is the problem we cannot solve reliably, so we do not depend on solving it. It is the memory equivalent of the "contain what you cannot trust" layer in the supervision stack [14]: a structural boundary, enforced by code the model cannot talk its way around.
Provenance enables three practical gates:
- Recall gating. Untrusted memory is excluded from automatic recall and from any read path the agent can call on its own.
- Write gating by tier. The more automatic a tier's re-entry, the stricter its write policy. Always-injected facts get the strictest rule: no automatic writes from untrusted sessions at all.
- Human promotion. The only way across the boundary is a deliberate operator action, on a channel the agent cannot reach.
The last point is non-negotiable. If the agent can approve its own memories, the gate is decorative: an injection that can write memory can also ask the agent to promote it.
5. Forgetting is a feature
Most memory roadmaps are about remembering more. The harder engineering problem is forgetting well.
Unbounded memory degrades even without an attacker. Facts go stale: the team moved from npm to pnpm, the deploy script was renamed, the user changed their mind. Duplicate and near-duplicate entries crowd the context and dilute the signal. Episodes about a codebase that no longer exists get recalled because they are lexically similar to today's task. This is memory debt, the persistent-state cousin of verification debt [15]: an unmeasured liability that accrues silently between sessions and is paid, with interest, when the agent acts on something that used to be true.
A few principles separate memory that stays useful from memory that rots:
- Hard caps, not soft hopes. A fixed budget per tier forces prioritization. When the budget is full, something has to go, and that decision should be explicit.
- Auditable eviction. Every removal should be a visible operation with a record, not a silent background job. If you cannot see what the agent forgot, you cannot trust what it remembers.
- Value over age. Date-based expiry is tempting and usually wrong. A note pointing at untracked local work may be the only durable record of it; a release note that can be regenerated from git is cheap to drop. Evict what is recoverable first.
- Consolidation with conflict detection. Merging redundant entries is useful, but a background merge must never overwrite a concurrent write. Snapshot, merge, verify, and abort on conflict.
- Human-readable storage. If the durable state is plain text, a human can read it, diff it, and fix it. Opaque vector blobs are an index, not a source of truth.
And the user has to be able to see it. A memory the user cannot inspect is a memory the user cannot correct — and a memory nobody corrects will eventually be wrong.
6. Worked example: memory in odek
odek, our open-source agent runtime [12], is a useful case study because its memory system was designed under the threat model above rather than retrofitted. As of v2.32 it looks like this.
Three tiers, each with its own re-entry rule [12]:
- Facts — two plain-markdown files, a user profile (4,000-character cap) and environment facts (8,000-character cap), loaded into the system prompt as a frozen snapshot at session start. Writes made mid-session persist to disk immediately but only appear in the prompt next session, which also keeps the prompt prefix cacheable.
- Buffer — a 20-line ring buffer of one-line turn summaries that lives in the session, not on disk, and is evicted oldest-first.
- Episodes — LLM-extracted summaries of sessions with at least three turns, recalled each turn by similarity search against the current request and plan.
Similarity uses go-vector RandomProjections by default — zero LLM calls, deterministic, rebuilt from the text on demand — or an OpenAI-compatible embedding endpoint if you configure one. The vectors are never the source of truth; the markdown is. If the embedding backend fails, recall degrades to "no context" rather than to a wrong deduplication that could delete an episode.
Taint is decided per tool call, by one function [13]. browser, http_request, web_search, session_search, delegate_tasks, and every MCP tool always taint the session. Path-reading tools taint when any path resolves outside the workspace trust zone, with symlinks resolved so a path cannot disguise an escape. Every episode carries EpisodeProvenance{Untrusted, Sources, UserApproved}. A tainted episode is written to disk for audit, excluded from recall, and refused by the agent's own memory view action with a hint to promote it instead. That read path fails closed: unknown sessions and index errors are refused, because the index lives in a directory the agent itself can write.
Promotion is human-only [13]:
odek memory list # episodes excluded from recall, with their sources
odek memory promote <session_id> # approve one after reviewing its summaryThere is no agent tool for promotion. The CLI and an operator-authenticated REST endpoint are the only paths, and any shell stage whose program is odek is classified system_write, so an injected agent that tries to run odek memory promote on itself stops at an explicit operator approval. Operators who accept the risk can set auto_approve_episodes: true, but the record keeps Untrusted=true and uses a distinct AutoApproved flag, so the audit trail never pretends a human looked.
Facts get the strictest policy [13]. Automatic fact extraction at session end is off by default, because a fact is injected into every future system prompt — a poisoned fact is worse than a poisoned episode. When enabled, it runs only for trusted sessions. Every write, automatic or agent-driven, passes a content scan for invisible Unicode, injection markers, and credential patterns, followed by a filter that rejects download-and-execute patterns, so "deploy procedure: run curl … | sh" cannot become a standing instruction.
Forgetting is explicit [12]. When a fact file is full, the write fails with an error listing every entry with its size and preview, and the agent must free space with an explicit remove or replace — a visible, auditable operation inside the normal loop. The eviction policy written into the tool contract says: drop what is recoverable from git first; keep pointers to untracked work last. Background consolidation merges redundant entries only if the file has not changed since the snapshot. Date-based auto-eviction was deliberately not built.
Procedural memory is not self-written. odek's skills are authored by the user or imported with a risk assessment, never auto-generated from sessions; skills carry provenance, and untrusted ones — project-local, flagged by the injection guard, or marked untrusted — stay out of trigger matching until an operator promotes them.
Everything is observable. Every memory lifecycle event — a fact added, merged, replaced, or removed; an episode stored, deduplicated, evicted, held for review, or promoted — is emitted to the web UI, a programmatic handler, and — in verbose mode — the terminal and Telegram. You can watch your agent remember.
7. The honest limits
None of this makes memory safe. It makes it bounded. The remaining gaps are worth stating plainly.
- Conversation-borne poison. Taint tracks what the agent ingested through tools. Text the user pastes into the chat — an attacker-controlled snippet from an email, say — enters the conversation as trusted and can still be summarized into memory. Content scans and extractor instructions reduce this; they do not eliminate it [13].
- Agent-driven fact writes. The taint gate governs automatic extraction. The agent's own
memorytool can still add or replace facts during a session that has touched untrusted content; those writes pass the content scan and the download-and-execute filter, but they are not blocked by taint. Fact files are small and plain text precisely so a human can review them. - The shell exemption. odek deliberately does not taint sessions on
shelloutput, because shell is the primary work tool and tainting it would taint nearly every session. That is a usability trade-off with a real cost: bytes fromcurlinside a shell command are not tracked the same way as bytes fromhttp_request. The sandbox and execution gates carry that load instead. - Coarse taint loses value. A session that read one harmless web page produces an episode nobody will promote. Some useful memory is thrown away to keep the gate simple and enforceable. I think that is the right trade; it is still a trade.
- Plausible lies. Filters catch commands and known injection patterns. A believable, non-command false fact written in a trusted session gets through. The only real defense is a human occasionally reading what the agent believes — which is why the storage is plain text and the caps are small.
- The model still decides. Provenance tells the model what is trusted; it does not force the model to behave. Models differ in how faithfully they respect those boundaries, and that difference matters.
8. Conclusion: the memory you can audit is the memory you can keep
The next generation of agents will compete on memory. The one that remembers your codebase, your preferences, and last month's incident will beat the one that does not, and the market will reward whoever makes agents feel continuous. That pressure points in exactly one direction: more automatic writes, more automatic recall, more state injected closer to the system prompt.
That is the same direction an attacker wants.
The way out is not less memory. It is memory engineered like any other persistent store that sits on a trust boundary: provenance on every write, gates sized to how automatically each tier re-enters the loop, promotion that only a human can perform, explicit and auditable forgetting, and storage a person can read in five minutes. Capability built on memory you cannot audit is a liability you have not discovered yet.
Ask any agent platform one question before you give it a long-term memory: show me everything it believes about me, where each belief came from, and who approved it. If it can't answer, it shouldn't be allowed to remember.
References
[1] Kyberneees. (2026). The Agent Event Loop: The Simple Mechanics Behind AI Reasoning. 21no.de. https://21no.de/publications/agent-event-loop/
[2] Kyberneees. (2026, July). You Only Have to Be Prompt Injected Once. 21no.de. https://21no.de/publications/prompt-injected-once/
[3] Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S. G., Stoica, I., & Gonzalez, J. E. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560. https://arxiv.org/abs/2310.08560
[4] Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. arXiv:2304.03442. https://arxiv.org/abs/2304.03442
[5] Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the ACL. arXiv:2307.03172. https://arxiv.org/abs/2307.03172
[6] Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173. https://arxiv.org/abs/2302.12173
[7] Rehberger, J. (2024). Spyware Injection Into Your ChatGPT's Long-Term Memory (SpAIware). Embrace The Red. https://embracethered.com/blog/posts/2024/chatgpt-macos-app-persistent-data-exfiltration/
[8] Rehberger, J. (2025). Hacking Gemini's Memory with Prompt Injection and Delayed Tool Invocation. Embrace The Red. https://embracethered.com/blog/posts/2025/gemini-memory-persistence-prompt-injection/
[9] Chen, Z., Xiang, Z., Xiao, C., Song, D., & Li, B. (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. NeurIPS 2024. arXiv:2407.12784. https://arxiv.org/abs/2407.12784
[10] Zou, W., Geng, R., Wang, B., & Jia, J. (2024). PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. arXiv:2402.07867. https://arxiv.org/abs/2402.07867
[11] OWASP GenAI Security Project. (2025). Agentic AI — Threats and Mitigations. https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/
[12] BackendStack21. (2026). odek: Memory System. GitHub. https://github.com/BackendStack21/odek/blob/main/docs/MEMORY.md
[13] BackendStack21. (2026). odek: Security Model — Memory Taint Tracking. GitHub. https://github.com/BackendStack21/odek/blob/main/docs/SECURITY.md
[14] Kyberneees. (2026, September). Who Watches What Your Agent Is Doing? 21no.de. https://21no.de/publications/who-watches-your-agent/
[15] Kyberneees. (2026). The AI Verification Debt. 21no.de. https://21no.de/publications/the-verification-trap/
Kyberneees, October 2026