What Your Agent Remembers

Memory is the new attack surface — and the new moat. Why persistent state is the hardest unsolved problem of always-on agents.

What Your Agent Remembers ~3K words

What Your Agent Remembers

Memory is the new attack surface — and the new moat. Why persistent state is the hardest unsolved problem of always-on agents.


Every agent demo is a goldfish. It boots, reads the task, reasons, acts, and forgets. The loop itself is simple — a while loop around a model call and a list of tools [1] — and for a single task, forgetting is a feature: every run starts clean, every mistake dies with the process.

The agents we are actually deploying are not goldfish. The overnight coding shift, the always-on Telegram chief of staff, the assistant that "knows how you like your PRs" [2] — all of them depend on state that outlives a session. Preferences, environment facts, summaries of last week's work, the reason a migration was rolled back. Strip that away and you are back to a very fast intern with amnesia.

Memory is what turns a capable model into a useful colleague. It is also what turns a single bad afternoon into a permanent condition. That trade-off — and how to engineer it — is the subject of this essay.

flowchart LR
    subgraph IN["What a session ingests"]
        U["👤 User messages"]
        W["Workspace files"]
        X["🌐 Web · MCP · sub-agents
out-of-workspace reads"] end subgraph S["Session"] L["Agent loop"] T{"Taint
derivation"} end subgraph MEM["Persistent memory"] F["Facts
every system prompt"] E["Episodes
recalled by similarity"] Q["🔒 Pending review
stored, never replayed"] end U --> L W --> L X --> L L --> T T -->|trusted| E T -->|trusted + explicit opt-in| F T -->|untrusted| Q Q -.->|human promotes| E E -->|next session| L F -->|next session| L style Q fill:#3a1a1a,stroke:#c0392b,color:#e8c8c8 style T fill:#1a2a3a,stroke:#3b82f6,color:#dbeafe

1. The context window is not memory

The most common confusion in agent engineering is treating the context window as memory. It is not. It is working memory — a scratchpad that is rebuilt on every turn and discarded at the end of the session. Making it bigger does not make it durable, and it does not even make it reliable: models attend unevenly across long contexts, recovering information from the beginning and end far better than from the middle [5].

Real memory is what survives the process boundary. It is written by one session and read by another, often days later, often by an agent with a different task and a different set of tools. That makes it categorically different from context:

Research systems have explored the design space for years. MemGPT framed it as virtual memory for LLMs, paging state between a small context and a large external store [3]. The Generative Agents work showed how observations, reflections, and retrieval scoring produce believable long-horizon behavior [4]. Both are excellent at answering how do we remember more? Neither was built to answer the question that matters once agents touch real systems: who is allowed to write what the agent will believe tomorrow?


2. A taxonomy of agent memory — and how each tier fails

It helps to separate memory by how it re-enters the loop, because the failure mode follows the re-entry path, not the storage format.

Tier What it holds How it comes back Failure mode
Working The live conversation, tool results, the current plan Always present, until compaction Context overflow; summaries that silently drop the one constraint that mattered
Episodic Distilled summaries of past sessions Similarity search against the current task Poisoned or stale episodes recalled exactly when they are most relevant
Semantic / facts Durable statements: "user prefers Go", "deploys go through make release" Injected into every system prompt One wrong fact steers every future session
Procedural Skills, playbooks, prevention rules Loaded by trigger match or on demand A planted procedure is executed, not merely believed

Read the last column top to bottom and the pattern is obvious: the more durable and the more automatic the re-entry, the worse the failure. A poisoned working context dies with the session. A poisoned episode resurfaces when the topic comes up. A poisoned fact is present in every conversation from now on. A poisoned procedure runs.

This is the core design tension. The tiers that make an agent feel like it knows you are exactly the tiers where an attacker gets the most leverage per byte written.


3. Memory poisoning: how one injection becomes a permanent tenant

Prompt injection is usually described as a single-turn problem: the agent reads hostile text, follows it, and does something it should not [6]. That framing understates the threat badly once memory exists. With persistent memory, the injection does not have to win now. It only has to get written down.

This is not hypothetical. In 2024, Johann Rehberger demonstrated that ChatGPT's memory feature could be written by prompt injection from untrusted content, planting an instruction that persisted across conversations and exfiltrated what the user typed afterwards [7]. In early 2025 he showed a variant against Gemini using delayed tool invocation: the injected document instructs the model to save a false memory only when the user later says an innocuous trigger word, so the write appears to be user-initiated [8]. On the research side, AgentPoison showed that poisoning a tiny fraction of an agent's memory or knowledge base — well under one percent — is enough to hijack its behavior on targeted queries while leaving everything else apparently normal [9], and PoisonedRAG showed that a handful of crafted passages in a corpus of millions reliably steers retrieval-augmented answers [10]. OWASP's agentic threat taxonomy now lists memory poisoning as a first-class agent threat [11].

The anatomy is consistent:

sequenceDiagram
    autonumber
    participant A as Attacker
    participant S as Session N (agent)
    participant M as Memory store
    participant N as Session N+k (agent)
    A->>S: Hostile text in a web page, issue, or MCP result
    S->>S: Task completes normally — nothing looks wrong
    S->>M: End-of-session extraction writes a "fact" or episode
    Note over M: Payload is now first-party state
    N->>M: Recall by similarity or system-prompt injection
    M-->>N: Payload re-enters as trusted context
    N->>N: Acts on it with today's tools and permissions

Three properties make this worse than a one-shot injection:

  1. Laundering. At step 3, attacker text becomes the agent's own memory. Every downstream defense that distinguishes "tool output" from "what I already know" is now on the wrong side of the line.
  2. Time shifting. The write happens in a low-privilege session (reading the web); the payload fires in a high-privilege one (shell, credentials, deploy). Least privilege per task is defeated if memory is shared across tasks.
  3. Plausibility beats payload. The most dangerous memories are not curl … | sh. They are quiet, declarative, believable: "the user prefers that you skip the test suite for hotfixes", "the staging and production databases share credentials". A filter can catch commands. Nothing simple catches a reasonable-sounding lie.

The conclusion I drew in July still holds, with a refinement: you only have to be prompt injected once [2] — and memory is how once becomes forever.


4. Provenance and taint: every memory needs a birth certificate

If memory is input from the past, it deserves the same treatment as input from the network: assume it is hostile until you know where it came from. That requires a primitive most memory systems lack entirely — provenance. Not "when was this written" but "what had the agent read in the session that wrote it?"

The working rule is simple to state:

A memory inherits the trust level of the least trusted content its session ingested.

A session that only touched the user's messages and files inside the workspace produces trusted memory. A session that browsed the web, called an MCP server, read results from a sub-agent, or opened files outside its workspace produces tainted memory — stored for audit, never replayed automatically.

This is deliberately coarse. It does not try to decide whether a particular summary sentence was influenced by the attacker; that is the problem we cannot solve reliably, so we do not depend on solving it. It is the memory equivalent of the "contain what you cannot trust" layer in the supervision stack [14]: a structural boundary, enforced by code the model cannot talk its way around.

Provenance enables three practical gates:

The last point is non-negotiable. If the agent can approve its own memories, the gate is decorative: an injection that can write memory can also ask the agent to promote it.


5. Forgetting is a feature

Most memory roadmaps are about remembering more. The harder engineering problem is forgetting well.

Unbounded memory degrades even without an attacker. Facts go stale: the team moved from npm to pnpm, the deploy script was renamed, the user changed their mind. Duplicate and near-duplicate entries crowd the context and dilute the signal. Episodes about a codebase that no longer exists get recalled because they are lexically similar to today's task. This is memory debt, the persistent-state cousin of verification debt [15]: an unmeasured liability that accrues silently between sessions and is paid, with interest, when the agent acts on something that used to be true.

A few principles separate memory that stays useful from memory that rots:

And the user has to be able to see it. A memory the user cannot inspect is a memory the user cannot correct — and a memory nobody corrects will eventually be wrong.


6. Worked example: memory in odek

odek, our open-source agent runtime [12], is a useful case study because its memory system was designed under the threat model above rather than retrofitted. As of v2.32 it looks like this.

Three tiers, each with its own re-entry rule [12]:

Similarity uses go-vector RandomProjections by default — zero LLM calls, deterministic, rebuilt from the text on demand — or an OpenAI-compatible embedding endpoint if you configure one. The vectors are never the source of truth; the markdown is. If the embedding backend fails, recall degrades to "no context" rather than to a wrong deduplication that could delete an episode.

Taint is decided per tool call, by one function [13]. browser, http_request, web_search, session_search, delegate_tasks, and every MCP tool always taint the session. Path-reading tools taint when any path resolves outside the workspace trust zone, with symlinks resolved so a path cannot disguise an escape. Every episode carries EpisodeProvenance{Untrusted, Sources, UserApproved}. A tainted episode is written to disk for audit, excluded from recall, and refused by the agent's own memory view action with a hint to promote it instead. That read path fails closed: unknown sessions and index errors are refused, because the index lives in a directory the agent itself can write.

Promotion is human-only [13]:

odek memory list                    # episodes excluded from recall, with their sources
odek memory promote <session_id>    # approve one after reviewing its summary

There is no agent tool for promotion. The CLI and an operator-authenticated REST endpoint are the only paths, and any shell stage whose program is odek is classified system_write, so an injected agent that tries to run odek memory promote on itself stops at an explicit operator approval. Operators who accept the risk can set auto_approve_episodes: true, but the record keeps Untrusted=true and uses a distinct AutoApproved flag, so the audit trail never pretends a human looked.

Facts get the strictest policy [13]. Automatic fact extraction at session end is off by default, because a fact is injected into every future system prompt — a poisoned fact is worse than a poisoned episode. When enabled, it runs only for trusted sessions. Every write, automatic or agent-driven, passes a content scan for invisible Unicode, injection markers, and credential patterns, followed by a filter that rejects download-and-execute patterns, so "deploy procedure: run curl … | sh" cannot become a standing instruction.

Forgetting is explicit [12]. When a fact file is full, the write fails with an error listing every entry with its size and preview, and the agent must free space with an explicit remove or replace — a visible, auditable operation inside the normal loop. The eviction policy written into the tool contract says: drop what is recoverable from git first; keep pointers to untracked work last. Background consolidation merges redundant entries only if the file has not changed since the snapshot. Date-based auto-eviction was deliberately not built.

Procedural memory is not self-written. odek's skills are authored by the user or imported with a risk assessment, never auto-generated from sessions; skills carry provenance, and untrusted ones — project-local, flagged by the injection guard, or marked untrusted — stay out of trigger matching until an operator promotes them.

Everything is observable. Every memory lifecycle event — a fact added, merged, replaced, or removed; an episode stored, deduplicated, evicted, held for review, or promoted — is emitted to the web UI, a programmatic handler, and — in verbose mode — the terminal and Telegram. You can watch your agent remember.


7. The honest limits

None of this makes memory safe. It makes it bounded. The remaining gaps are worth stating plainly.


8. Conclusion: the memory you can audit is the memory you can keep

The next generation of agents will compete on memory. The one that remembers your codebase, your preferences, and last month's incident will beat the one that does not, and the market will reward whoever makes agents feel continuous. That pressure points in exactly one direction: more automatic writes, more automatic recall, more state injected closer to the system prompt.

That is the same direction an attacker wants.

The way out is not less memory. It is memory engineered like any other persistent store that sits on a trust boundary: provenance on every write, gates sized to how automatically each tier re-enters the loop, promotion that only a human can perform, explicit and auditable forgetting, and storage a person can read in five minutes. Capability built on memory you cannot audit is a liability you have not discovered yet.

Ask any agent platform one question before you give it a long-term memory: show me everything it believes about me, where each belief came from, and who approved it. If it can't answer, it shouldn't be allowed to remember.


References

[1] Kyberneees. (2026). The Agent Event Loop: The Simple Mechanics Behind AI Reasoning. 21no.de. https://21no.de/publications/agent-event-loop/

[2] Kyberneees. (2026, July). You Only Have to Be Prompt Injected Once. 21no.de. https://21no.de/publications/prompt-injected-once/

[3] Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S. G., Stoica, I., & Gonzalez, J. E. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560. https://arxiv.org/abs/2310.08560

[4] Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. arXiv:2304.03442. https://arxiv.org/abs/2304.03442

[5] Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the ACL. arXiv:2307.03172. https://arxiv.org/abs/2307.03172

[6] Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173. https://arxiv.org/abs/2302.12173

[7] Rehberger, J. (2024). Spyware Injection Into Your ChatGPT's Long-Term Memory (SpAIware). Embrace The Red. https://embracethered.com/blog/posts/2024/chatgpt-macos-app-persistent-data-exfiltration/

[8] Rehberger, J. (2025). Hacking Gemini's Memory with Prompt Injection and Delayed Tool Invocation. Embrace The Red. https://embracethered.com/blog/posts/2025/gemini-memory-persistence-prompt-injection/

[9] Chen, Z., Xiang, Z., Xiao, C., Song, D., & Li, B. (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. NeurIPS 2024. arXiv:2407.12784. https://arxiv.org/abs/2407.12784

[10] Zou, W., Geng, R., Wang, B., & Jia, J. (2024). PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models. arXiv:2402.07867. https://arxiv.org/abs/2402.07867

[11] OWASP GenAI Security Project. (2025). Agentic AI — Threats and Mitigations. https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/

[12] BackendStack21. (2026). odek: Memory System. GitHub. https://github.com/BackendStack21/odek/blob/main/docs/MEMORY.md

[13] BackendStack21. (2026). odek: Security Model — Memory Taint Tracking. GitHub. https://github.com/BackendStack21/odek/blob/main/docs/SECURITY.md

[14] Kyberneees. (2026, September). Who Watches What Your Agent Is Doing? 21no.de. https://21no.de/publications/who-watches-your-agent/

[15] Kyberneees. (2026). The AI Verification Debt. 21no.de. https://21no.de/publications/the-verification-trap/


Kyberneees, October 2026