Short-Term vs Long-Term Memory in AI Agents

Short-Term vs Long-Term Memory in AI Agents

Z
Zahid Hasan Tonmoy
September 29, 2026
10 min read
🇧🇩 বাংলায় পড়ুন
|
0 views
Audio Overview & Podcast
~1 min quick overview
0%

Last post was about picking a language and the handful of tools an agent actually needs. Whichever you picked, the first real problem shows up once the agent runs past a few turns: the conversation keeps growing, every call gets slower and more expensive, and the agent still forgets something you told it last week, in a different session entirely. Those are two different problems wearing the same name.

Short-term memory is the running conversation — everything said in this session, sent back to the model on every call so it has context. Long-term memory is anything saved on purpose so it survives after the session ends and can be pulled back in later, in a different conversation entirely. The two need completely different handling: one has to stay bounded in size, the other has to be searchable.

Short-Term Memory Is Just the Array You Already Know

The post on an agent's four parts called this out first: short-term memory is the messages array sent with every call. Nothing mysterious about it — it's the literal conversation history, and the model reads all of it, every time, because it remembers nothing between calls on its own.

That's exactly what breaks it at scale. Every message ever sent in the session rides along on every future call: cost climbs with the conversation's length, not just the current question, and a long enough history starts pushing against the model's context window. There's also a subtler cost — a very long history can bury a detail from ten turns ago in the middle of a wall of text the model has to weigh against everything else.

It's close to how you'd hold a phone number in your head for as long as the call lasts, then forget it the moment you hang up unless you actually write it down somewhere. The model's short-term memory works the same way: it exists for exactly as long as the array does, and the instant that array is gone or trimmed, so is everything in it, whether or not it mattered.

Long-Term Memory Is a Different Shape Entirely

Long-term memory isn't a bigger version of the message array. It's a separate store — a file, a database row, a vector store — that the agent writes to on purpose and reads from selectively, not by replaying the whole thing. Compare that to the RAG post: retrieval augmented generation pulls relevant passages from documents someone else wrote. Long-term memory does the identical retrieval, aimed at facts the agent or the user decided were worth keeping. Once it's more than a handful of notes, agent memory basically is RAG, just pointed inward instead of at a document set.

Both, in One File

📊 Architecture DiagramArchitecture Flow
Interactive diagram rendering...
flowchart TD
    U[New user message] --> ST[Short-term: this conversation's message array]
    ST -->|too long| C[Summarize older turns, keep recent ones]
    C --> ST
    U --> LT[Long-term: facts saved from any past session]
    LT -->|relevant to this message| R[Add matching facts to the system prompt]
    ST --> M[Model]
    R --> M
System architecture specification and node flow: flowchart TD U[New user message] --> ST[Short-term: this conversation's message array] ST -->|too long| C[Summarize older turns, keep recent ones] C --> ST U --> LT[Long-term: facts saved from any past session] LT -->|relevant to this message| R[Add matching facts to the system prompt] ST --> M[Model] R --> M

Flow summary: A new message checks both memories at once. The short-term array gets compacted if it has grown too long, and long-term storage is searched for anything relevant, with matches added to the system prompt; both then reach the model together.

Install the SDK with npm install @anthropic-ai/sdk, set ANTHROPIC_API_KEY, save this as agent/memory-agent.ts and run npx tsx agent/memory-agent.ts.

ts
// agent/memory-agent.ts
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

type Memory = { fact: string; vector: number[] };
const longTerm: Memory[] = [];

function tokenize(text: string): string[] {
  return text.toLowerCase().match(/[a-z0-9]+/g) ?? [];
}

function vectorize(text: string, vocabulary: string[]): number[] {
  const counts = new Map<string, number>();
  for (const token of tokenize(text)) counts.set(token, (counts.get(token) ?? 0) + 1);
  return vocabulary.map((word) => counts.get(word) ?? 0);
}

function cosineSimilarity(a: number[], b: number[]): number {
  let dot = 0;
  let normA = 0;
  let normB = 0;
  for (let i = 0; i < a.length; i++) {
    dot += a[i] * b[i];
    normA += a[i] * a[i];
    normB += b[i] * b[i];
  }
  if (normA === 0 || normB === 0) return 0;
  return dot / (Math.sqrt(normA) * Math.sqrt(normB));
}

function rememberFact(fact: string): void {
  const vocabulary = Array.from(new Set([fact, ...longTerm.map((m) => m.fact)].flatMap(tokenize)));
  longTerm.push({ fact, vector: vectorize(fact, vocabulary) });
}

function recallRelevant(query: string, minScore = 0.2): string[] {
  if (longTerm.length === 0) return [];
  const vocabulary = Array.from(new Set([query, ...longTerm.map((m) => m.fact)].flatMap(tokenize)));
  const queryVector = vectorize(query, vocabulary);
  return longTerm
    .map((m) => ({ fact: m.fact, score: cosineSimilarity(queryVector, vectorize(m.fact, vocabulary)) }))
    .filter((m) => m.score >= minScore)
    .sort((a, b) => b.score - a.score)
    .slice(0, 2)
    .map((m) => m.fact);
}

const tools: Anthropic.Tool[] = [
  {
    name: 'remember',
    description: 'Save a fact about the user that should be recalled in future sessions, such as a preference or a standing constraint.',
    input_schema: {
      type: 'object',
      properties: { fact: { type: 'string' } },
      required: ['fact'],
    },
  },
];

const messages: Anthropic.MessageParam[] = [];
const turnBoundaries: number[] = [];
const KEEP_RECENT_TURNS = 2;
const COMPACT_AFTER_TURNS = 4;

async function compactIfNeeded(): Promise<void> {
  if (turnBoundaries.length <= COMPACT_AFTER_TURNS) return;

  const cutBoundary = turnBoundaries[turnBoundaries.length - KEEP_RECENT_TURNS];
  const older = messages.slice(0, cutBoundary);
  const recent = messages.slice(cutBoundary);

  const summaryResponse = await client.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 200,
    system: 'Summarize this conversation in 2-3 sentences, keeping only details that matter for future turns.',
    messages: older,
  });

  const summaryBlock = summaryResponse.content.find((b): b is Anthropic.TextBlock => b.type === 'text');

  messages.length = 0;
  messages.push({ role: 'user', content: `Earlier in this conversation: ${summaryBlock?.text ?? ''}` });
  messages.push({ role: 'assistant', content: 'Got it, continuing from there.' });
  messages.push(...recent);

  const shift = 2 - cutBoundary;
  const kept = turnBoundaries.slice(turnBoundaries.length - KEEP_RECENT_TURNS);
  turnBoundaries.length = 0;
  turnBoundaries.push(...kept.map((i) => i + shift));
}

async function runAgent(userText: string): Promise<string> {
  const relevant = recallRelevant(userText);

  turnBoundaries.push(messages.length);
  messages.push({ role: 'user', content: userText });
  await compactIfNeeded();

  for (let step = 0; step < 4; step++) {
    const response = await client.messages.create({
      model: 'claude-sonnet-5',
      max_tokens: 500,
      ...(relevant.length > 0 ? { system: `Known about the user from earlier sessions: ${relevant.join('; ')}` } : {}),
      tools,
      messages,
    });

    messages.push({ role: 'assistant', content: response.content });

    if (response.stop_reason !== 'tool_use') {
      const block = response.content.find((b): b is Anthropic.TextBlock => b.type === 'text');
      return block?.text ?? '';
    }

    const results: Anthropic.ToolResultBlockParam[] = [];
    for (const block of response.content) {
      if (block.type === 'tool_use') {
        if (block.name === 'remember') {
          rememberFact((block.input as { fact: string }).fact);
        }
        results.push({ type: 'tool_result', tool_use_id: block.id, content: JSON.stringify({ ok: true }) });
      }
    }
    messages.push({ role: 'user', content: results });
  }

  return 'Stopped: step limit reached';
}

runAgent('I only drink green tea, never coffee.').then(console.log);

recallRelevant reuses the same tokenize-and-cosine-similarity retriever from the RAG post, pointed at longTerm instead of a document array. compactIfNeeded is the short-term side: once the conversation passes four turns, it summarizes everything except the most recent two and replaces the older messages with a single synthetic exchange, tracking exact turn boundaries in turnBoundaries so a compaction can never cut a tool call away from its result — a real risk once a turn contains more than a plain question and answer, since the model's API rejects a tool call left without a matching result.

What I Actually Found Testing This

Before writing this up, I ran the file through several turns: one where the agent saves a preference with the remember tool, a few unrelated exchanges to force a compaction, then a later turn asking about that same preference. Compaction worked cleanly. Recall didn't.

The saved fact was drinks only green tea, never coffee. The later question was What do I usually drink?. Zero word overlap between them — drink in the question, drinks in the fact, and a plain word-count vector treats those as two entirely different tokens. The similarity score came back as exactly zero, well under the cutoff, so nothing got added to the system prompt and the agent had no idea it had ever been told this.

This is the same limitation the RAG post ran into, and it's worse here: a document usually repeats its own key terms somewhere, but a memory gets saved once in whatever words came up in that moment, and asked about later in completely different ones. A toy retriever built for demonstrating the mechanism runs into this fast. Fixing it for real means what the RAG post already pointed at — an actual embedding model that captures meaning, not just shared characters, since Anthropic's own docs recommend Voyage AI for exactly this because Anthropic doesn't ship one itself.

Common Mistakes With Agent Memory

  • Treating everything as worth remembering long-term. Most of a conversation is only useful within that conversation. Save facts that would still matter weeks later — preferences, constraints, decisions — not the details of the current exchange.
  • No deduplication on save. In the same test run, the agent called remember with the identical fact twice across two different turns, and the code happily stored it both times. A real system should check for a near-duplicate before writing a new entry.
  • Compacting at an arbitrary message count instead of a turn boundary. Slicing by a fixed number of messages can land in the middle of a tool call and its result, which the API will reject outright. Track turn boundaries explicitly, the way turnBoundaries does here.
  • Assuming a summary preserves everything worth keeping. It won't, by design — that's the whole point of compacting. Anything that truly can't be lost belongs in long-term memory as an explicit fact, not hoped for inside a summary of an old conversation.

What's Next

Memory tells an agent what it knows. Getting it to act on that reliably, across several steps toward one goal instead of reacting one call at a time, is worth a post of its own.

Frequently Asked Questions

Is long-term memory basically the same thing as RAG?

Mechanically, yes, once it's more than a handful of notes. Both embed some text, compare it against a stored set, and pull back the closest matches. RAG points that at documents someone else wrote; long-term memory points the identical mechanism at facts the agent or the user decided to save.

How do I decide what's worth saving to long-term memory?

Not every message — most of a conversation only matters within that conversation, which is exactly what short-term memory already covers for its length. Save what would still be true and useful weeks later: a stated preference, a standing constraint, a decision that shouldn't need re-explaining next time.

Won't compacting the conversation lose important details?

Some, yes, and that's an intentional trade-off, not a bug. Keep the most recent turns verbatim so nearby context stays exact, and summarize only what's older. Anything that truly must never be lost belongs in long-term memory as an explicit saved fact, not left to survive by chance inside a summary.

img2
img2

img1
img1

thumbnail
thumbnail

🟢 Available for Freelance & Contract Work

Need a High-Performance Web App or Custom AI Solution?

I help founders, businesses, and engineering teams build lightning-fast web applications, resilient backend architectures, and intelligent AI workflows. Have an idea in mind? Let’s bring it to life.

Fast MVP Launch (2–4 weeks)
AI Agent & LLM API Integration
100/100 Core Web Vitals & SEO
Production-Ready Clean Architecture

Found this article helpful?

Give some claps to support more in-depth engineering logs!

0 views
Z

Zahid Hasan Tonmoy

Author & Developer

MERN Full Stack Developer & AI Agent Developer based in Dhaka, Bangladesh. Writing about web development, React, PostgreSQL and my learning journey.

Related Articles

Stay Updated

Subscribe to get insights on full-stack architecture, AI agent engineering, and web development.