When I first deployed the support-automation system, retrieval worked. The agent could find the right documentation, surface the relevant ticket history, and generate a coherent response. But within weeks, I started noticing a pattern in the feedback: users were frustrated that the system kept "forgetting" things.
A customer would contact support, we'd resolve a complex billing issue, and two weeks later they'd return with a related problem. The agent had no memory of the previous interaction. It would start from scratch — re-asking questions already answered, suggesting steps already tried. The technical retrieval was fine. The problem was that retrieval and memory are not the same thing.
That gap sent me down a deep rabbit hole of agent memory architecture. Here's what I built, what broke, and what actually works.
The Stateless Trap
Most AI applications are stateless by design. Every request comes in cold — the model has no knowledge of previous sessions unless you explicitly provide it in the context window.
For a chatbot answering FAQs, that's fine. For an agent handling ongoing customer relationships or complex multi-session workflows, it's a critical failure mode.
The stateless trap has three stages:
Stage 1 — It works in demos. Short, contained conversations work perfectly. The context window covers everything needed.
Stage 2 — It fails at scale. Real users have history. They expect continuity. When the agent treats every interaction as session zero, trust erodes fast.
Stage 3 — The workaround is worse. Naive fixes — like stuffing all history into the context window — hit token limits, slow down inference, and introduce noise that degrades response quality.
Memory architecture is the real solution. But memory is not a single thing.
Four Types of Memory That Matter
After building and iterating on the support-automation system, I settled on a taxonomy of four memory types. Each serves a different role:
Short-Term Context
This is what lives in the current context window. The conversation so far, the tools called, the documents retrieved. It's fast, zero latency, and sufficient for most single-session tasks.
The mistake is trying to use short-term context for everything. Once you need cross-session continuity, you've already failed if this is your only memory layer.
Episodic Memory
Episodic memory stores specific past events — individual interactions, decisions made, outcomes observed. In the support-automation system, this meant recording each resolved ticket: what the problem was, what we tried, what worked.
When a user returned with a related issue, the agent could retrieve the relevant episode and say: "Last time we saw this pattern, the root cause was in the billing processor's retry logic. Let me check that first."
The key design decision: store episodes as structured summaries, not raw conversation logs. Raw logs are noisy and expensive to retrieve from. A well-structured episode captures the signal — the problem classification, the resolution path, the outcome — in a form the agent can actually use.
Semantic Memory
Semantic memory is general knowledge — facts, relationships, patterns that aren't tied to a specific episode. In the support-automation system, this became the accumulated understanding of which error patterns indicated which root causes.
Technically, I implemented this as a vector store: every resolved ticket generated embeddings, and over time the system built up a dense semantic map of the problem space. New issues could be matched against historical patterns using similarity search.
The difference from episodic memory: semantic memory compresses and consolidates. You don't need to remember every specific ticket — you need to remember the pattern "error code 4XX + billing_processor tag → almost always retry logic." That's semantic knowledge, and it gets richer the longer the system runs.
For embeddings, I paired this with a dedicated embeddings model (the semantic storage layer doesn't need to match the LLM doing the reasoning) and stored vectors in pgvector on the same Postgres instance handling the rest of the system state.
Procedural Memory
Procedural memory is the agent's learned skills — not what happened, but how to do things. In the Python orchestration framework, this shows up as the set of tools and workflows the agent has available. As the system matures, the "procedures" improve: better prompts, refined tool signatures, updated decision trees.
This is the memory type that evolves slowest but has the highest leverage. A better diagnostic procedure means every future session benefits, not just the next one.
The Memory Stack in Practice
Here's the architecture that held up in production:
The retrieval step is where most implementations go wrong. A common mistake is running all memory queries in sequence. In the support-automation system, I run episodic and semantic retrieval in parallel, then merge the results before passing them to the model.
What Memory Actually Changed
After deploying the full memory stack, two things changed noticeably:
Resolution time dropped. The agent stopped re-diagnosing problems it had already seen. Returning users with recurring issues got faster responses because the agent could immediately surface what had worked before.
User trust increased. This one surprised me. Users didn't know the technical details of what changed. They just noticed that the system remembered them. That felt intelligent in a way that raw accuracy never quite did.
The Memory Consolidation Problem
There's a challenge that doesn't show up until you've been running memory for a few months: memory bloat. Raw episodic storage grows without bound. Older episodes become noise. The semantic store drifts as patterns evolve.
I handle this with periodic consolidation:
- Episodic pruning: Archive episodes older than 90 days. For repeat users, keep a rolling summary of all prior sessions rather than individual episode records.
- Semantic refresh: Re-index the semantic store quarterly. New patterns have emerged; old ones may no longer be representative.
- Weight recent context: In the retrieval step, weight recent episodes more heavily than old ones.
This is still the least mature part of the system. Memory consolidation is genuinely hard. The key insight: treat memory maintenance as a first-class operational concern, not an afterthought.
The Knowledge Graph Layer
For the Company Brain — the cross-project documentation and context system I use in my own work — I extended the memory architecture with an explicit knowledge graph layer. Rather than treating each document or note as an independent entity, the knowledge graph models the relationships between them.
An agent working in this system can traverse relationships: "This decision connects to this project, which depends on this external service, which has these known failure modes." That kind of contextual traversal isn't possible with pure vector similarity — it requires explicit relationship modeling.
For the support-automation system, I haven't added graph structure yet. The episodic and semantic combination handles the volume and query patterns well enough. But for knowledge-dense domains where relationships between entities matter, the graph layer is worth the engineering complexity.
Practical Starting Point
If you're adding memory to an existing agent, don't try to build all four layers at once. The order of return:
-
Start with episodic memory — store structured summaries of completed sessions and retrieve the most recent few for returning users. This gets you most of the "it remembers me" benefit with the least complexity.
-
Add semantic retrieval once your episodic store is large enough to be noisy. Embeddings let you find relevant patterns without having to retrieve every historical episode.
-
Revisit procedural memory periodically — your tools and workflows will improve naturally as you see where agents struggle; build that improvement into the procedure, not just the data.
-
Design consolidation early, even if you don't run it right away. The schema decisions you make for episodic storage determine how easily you can consolidate later.
Memory is what turns a capable AI tool into an agent that feels intelligent to the people using it. The technical lift is real, but the behavioral shift in how users perceive and trust the system is worth it.
The agent-memory-systems learning path on this site goes deeper on each of these patterns — from implementing embeddings to handling production memory consolidation. If you're building a stateful agent, that's a good next step.