At a glance
Key takeaways
- The model is stateless; memory is what your code puts back into the prompt.
- Short-term: keep recent turns verbatim and fold older ones into a summary, within a token budget.
- Long-term: episodic (events), semantic (facts), procedural (how-to), recalled by similarity.
- Write selectively, never store secrets, supersede rather than overwrite, and keep provenance.
- Isolate by partitioning on tenant and user, supplied by authentication, never by the query.
- Keep task state in a database; the model proposes, code validates and writes.
Level 2
How it works, from scratch
What follows builds the four pieces, each small enough to read in one sitting, and shows what each one prevents.
A language model remembers nothing between calls. Every call starts from a blank page plus whatever you put in the prompt. "Memory" is therefore entirely your system: what you keep, where you keep it, what you put back into the prompt, and who is allowed to see it. This lesson builds four pieces: short-term memory, long-term memory with three kinds of record, the safety rules around it (write policy, conflicts, isolation, deletion), and task state kept in a database instead of in the conversation.
Figure 1 · Diagram
flowchart LR
subgraph Prompt["What the model sees on this call"]
S[System prompt<br/>+ summary of older turns]
R[Recalled long-term memories]
M[Recent messages, word for word]
end
STM[(Short-term memory<br/>this conversation)] --> S
STM --> M
LTM[(Long-term memory<br/>per tenant, per user)] -->|similarity search| R
DB[(Task state<br/>SQLite)] -->|current step| S
Model[Model] -->|proposes updates| Code[Your code]
Code -->|validated writes| LTM
Code -->|validated writes| DB
Chapter 1
Short-term memory: the conversation, inside a budget
Everyday picture A flip chart in a long meeting. When the page fills up, you don't find a bigger pad. You tear off the old pages and start a fresh one with one line at the top: "Earlier: agreed on the budget, rejected vendor B."
Tiny worked example Six question-and-answer exchanges about invoices, 17 or 18 tokens per message (210 tokens in all), and a budget of 110 tokens. The last four messages stay word for word: 18 + 17 + 18 + 17 = 70 tokens, which leaves 110 - 70 = 40 tokens for the summary. The summary keeps the first sentence of each older message, but all eight of those would take 65 tokens, so the oldest drop out, one at a time, until the rest fit in 40:
summary : Earlier in this conversation: Question 3 is about invoices; Answer 3 lists the invoice;
Question 4 is about invoices; Answer 4 lists the invoice (36 tokens)
messages : Question 5 ..., Answer 5 ..., Question 6 ..., Answer 6 ... (verbatim, 70 tokens)
Total sent: 36 + 70 = 106 tokens, under the 110 budget. Exchanges 1 and 2 are gone from the prompt entirely. That's the price of a fixed budget, and it's why anything worth keeping forever belongs in long-term memory (section 2), not in the conversation.
Figure 2 · Chart
Sending the whole history climbs without end, while the managed history levels off just under its 300-token budget
The code ShortTermMemory.context() returns (summary, messages). The
summary goes in the system prompt rather than as a message, so user and
assistant turns still alternate as the API requires.
In code: ShortTermMemory.add appends a message and
ShortTermMemory.tokens totals the history; first_sentences is the
default summarizer, keeping the first sentence of each folded message.
ShortTermMemory.context gives the summary only the budget the recent messages leave
and drops the oldest folded messages until it fits. A production system
often re-summarizes the summary with an LLM call instead, trading an
extra call for losing less.
Why it matters Context rot, where quality drops as a session grows, is one of the most common agent failures. Summarize, trim, or reset with a written hand-off.
Chapter 2
Long-term memory: three kinds of record
Everyday picture Three notebooks: a diary of what happened (episodic: "on 18 Sept Alice rejected Globex"), an encyclopedia of facts (semantic: "Alice's fiscal year starts in April"), and a habit, the way you've learned to do something (procedural: "to reconcile, match each invoice to its payment, then list mismatches").
Tiny worked example Alice has one memory of each kind. The question "what
happened with Globex?" is embedded and compared with each memory (cosine
similarity, see primer.ml.embeddings.similarity), and the diary entry
comes back first. Asking with kinds=("procedural",) searches only habits.
| kind | stored as | typical recall trigger |
|---|---|---|
| episodic | dated event | "what happened with...", "last time..." |
| semantic | keyed fact (fiscal_year_start) |
any question the fact answers |
| procedural | how-to steps | "how do I...", before starting a known task |
Figure 3 · Chart
Each question recalls the right kind: the Globex question scores the diary entry highest, the how-to question the procedure
primer.agents.rag) over the agent's own history.In code: LongTermMemory.remember stores a MemoryRecord of one kind
with its embedding and reports back in a WriteResult.
LongTermMemory.recall returns the current records in a scope most similar
to the query, optionally limited to some kinds.
Chapter 3
The hard parts: what to write, conflicts, isolation, deletion
What to write. Everyday picture: a good assistant's notebook has
"prefers morning meetings", not "said thanks at 3:02pm", and never your
bank PIN. worth_remembering skips small talk and refuses anything that
looks like a secret, since memory is read back into prompts and shown in
exports.
Conflicts. Worked example: Alice said her fiscal year starts in April,
then corrected it to July. Both are stored under the key fiscal_year_start.
The April record is marked superseded by the July one, recall returns only
July, and history() still shows both with their source, so "why did the agent
think April?" has an answer.
Isolation. Everyday picture: separate locked filing cabinets per
company, not one cabinet with a "please only read your own folder" sign.
Multi-tenant means one system serves many customers (tenants). Every
memory call names a Scope (tenant + user), and storage is partitioned
by it, so a lookup starts inside the caller's own cabinet and can't reach
another.
Figure 4 · Diagram
flowchart TD Q[recall scope=globex/carol, query=...] --> P[Open partition<br/>tenant=globex, user=carol] P --> S[Similarity search<br/>inside this partition only] S --> R[Results] A[(acme/alice partition)] -.-x S
Figure 5 · Chart
Acme's secret scores a perfect match to Carol's quoted query, yet it sits outside her partition and is never searched
Deletion. Users may need to see, correct or delete what's remembered,
and privacy law such as the GDPR (the EU's General Data Protection
Regulation) can require it. export(scope) shows everything including
superseded facts, and delete_user(scope) erases one user without touching
colleagues.
In code: LongTermMemory.remember applies the write policy and marks
an older record with the same key as superseded; LongTermMemory.history
lists every value a key has had. LongTermMemory keeps one partition per
Scope, and LongTermMemory.export, LongTermMemory.forget and
LongTermMemory.delete_user show, remove one record, and erase a user.
Chapter 4
Task state outside the model
Everyday picture A checklist on a clipboard. The assistant can suggest ticking a box, but only the supervisor holds the pen. If the assistant goes home sick, the next person picks up the clipboard and starts at the first unticked box.
Tiny worked example
steps : fetch_invoices, fetch_payments, match
model proposes : {"step": "fetch_invoices", "status": "done", "output": "4 invoices"} -> written
model proposes : {"step": "match", "status": "done"} -> refused: the current step is fetch_payments
process crashes after fetch_payments is done
new process, same file : next_step() -> "match"
Figure 6 · Diagram
sequenceDiagram
participant M as Model
participant C as Your code
participant D as SQLite
M->>C: propose {"step": "fetch_invoices", "status": "done"}
C->>D: read current step
D-->>C: fetch_invoices
C->>D: write status=done (transaction)
M->>C: propose {"step": "match", "status": "done"}
C-->>M: refused: current step is fetch_payments
Note over C,D: crash, restart
C->>D: next_step?
D-->>C: match
In code: TaskStateStore keeps tasks and steps in SQLite.
TaskStateStore.create_task writes a task and its steps in one transaction,
TaskStateStore.next_step finds the first unfinished step, and
TaskStateStore.apply validates a proposal, raising InvalidUpdate for a
bad status or a skipped step, before writing it.
Why it matters Authoritative state in a database makes runs inspectable, resumable and auditable. That's much of the difference between a demo and a production system.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: How would you design long-term memory for a multi-tenant agent platform?Think it through, then reveal
A: Every read and write takes a scope (tenant, user) from the authenticated session, and storage is partitioned by it: separate namespaces, collections or row-level security, never a global index filtered after the search. Store typed records (episodic, semantic with keys, procedural) with embeddings, source and time. Apply a write policy (durable, non-secret), supersede conflicting facts instead of overwriting, retrieve the top few by similarity within scope, and support export and deletion per user. Test isolation with adversarial queries.
Question 2Q: A long support chat gets worse and more expensive over time. Why, and what helps?Think it through, then reveal
A: Every call re-sends the whole history, so cost rises, and models use details in the middle of long contexts less reliably. Keep a budget: recent turns verbatim, older ones summarized, important facts promoted to long-term memory, or reset with a written hand-off.
Question 3Q: The user corrects a fact the agent remembered. What should happen?Think it through, then reveal
A: Store the new value under the same key, mark the old one as superseded by the new (don't delete it silently), and recall only current records. Keep the source of each so the change is explainable.
Question 4Q: Why keep task state in a database when the model can "remember" progress in the conversation?Think it through, then reveal
A: The conversation is lost on a crash, and it can be summarized, trimmed or misread. A database is authoritative, survives restarts, can be queried by people and monitoring, and lets code enforce rules (no skipping steps) that a prompt can only request.
Primary sources
The papers behind this lesson
It introduced a memory stream of observations retrieved by relevance, recency and importance, plus periodic reflection into higher-level memories, a template for long-term agent memory.
The paper ↗It treats the context window like RAM and external storage like disk, with the agent paging information in and out, which is the short-term/long-term split made explicit.
The paper ↗Researcher's shelf
Further reading
- Lilian Weng, LLM Powered Autonomous Agents (memory section): https://lilianweng.github.io/posts/2023-06-23-agent/
- LangGraph docs (short- and long-term memory, persistence): https://langchain-ai.github.io/langgraph/
- SQLite, transactions: https://www.sqlite.org/lang_transaction.html
- GDPR, right to erasure (Art. 17): https://gdpr-info.eu/art-17-gdpr/
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.