What Is RAG? How AI Agents Retrieve and Use Real Knowledge
Learning Journal

What Is RAG? How AI Agents Retrieve and Use Real Knowledge

Z
Zahid Hasan Tonmoy
September 30, 2026
9 min read
🇧🇩 বাংলায় পড়ুন
|
0 views
Audio Overview & Podcast
~1 min quick overview
0%

The last post covered tool calling — how an agent reaches out and acts on the world through a function. This one is about a quieter problem: an agent can have every tool it needs and still answer wrong, simply because it doesn't know something. Not because it's incapable of knowing it — because that fact was never in its training data, or it changed after training ended, or it lives in a document only your company has.

Ask a model a question about a policy your team wrote last week, and it will either say it doesn't know or, worse, guess with total confidence. Retraining the model every time a document changes isn't realistic. Retrieval Augmented Generation is the answer that doesn't require retraining at all.

What Is RAG?

Retrieval Augmented Generation, or RAG, is a technique where relevant text is pulled from a knowledge source at the moment a question is asked and inserted into the model's prompt before it answers. Instead of relying only on what the model learned during training, the model reads a handful of retrieved passages alongside the question and answers from those. Update the underlying documents and the next answer reflects the change immediately, with no retraining involved.

Why an Agent Needs This

A model's knowledge is frozen at the point it was trained, and even within that, it can't possibly memorize every internal document, changelog or pricing sheet a specific company keeps. Two options exist for closing that gap. Fine-tuning bakes new information into the model's weights through additional training — expensive, slow, and awkward every time a source document changes. RAG leaves the model untouched and instead fetches the relevant text at request time. For information that changes on any kind of regular schedule — docs, policies, prices, code comments — RAG is almost always the cheaper, faster-to-update option.

The Three Moving Pieces

A RAG pipeline has three jobs, and it helps to think of them as separate from each other even when the code lives in one file.

Embedding. Turn a piece of text into a vector — a list of numbers positioned so that texts with similar meaning end up near each other in that number space. Both your documents and the incoming question go through this step.

Retrieval. Compare the question's vector against every stored document vector using a similarity measure, most commonly cosine similarity, and keep only the closest matches.

Augmentation. Insert those matches into the prompt as context, then let the model generate its answer from the question plus that context, instead of from the question alone.

📊 Architecture DiagramArchitecture Flow
Interactive diagram rendering...
flowchart LR
    Q[User question] --> E[Turn the question into a vector]
    E --> S[Compare against every stored document vector]
    S --> K[Keep only the closest matches above a threshold]
    K --> P[Add those chunks to the prompt as context]
    P --> M[Model generates the answer]
System architecture specification and node flow: flowchart LR Q[User question] --> E[Turn the question into a vector] E --> S[Compare against every stored document vector] S --> K[Keep only the closest matches above a threshold] K --> P[Add those chunks to the prompt as context] P --> M[Model generates the answer]

Flow summary: The question becomes a vector, gets compared against every stored document's vector, and only the closest matches above a threshold get added to the prompt as context before the model writes its answer.

A Working Example You Can Run Today

Anthropic doesn't ship its own embedding model — its own documentation points developers to Voyage AI as the recommended provider instead. That's the right call for anything real, but it means a genuine end-to-end example needs a second API key just to demonstrate retrieval. So the version below swaps a real embedding model for a small hand-written one: word counts turned into vectors, compared with cosine similarity. It's not how you'd retrieve at scale, but it's the same three steps, and the whole thing runs with nothing but an Anthropic key.

Install the SDK with npm install @anthropic-ai/sdk, set ANTHROPIC_API_KEY, save this as agent/simple-rag.ts and run npx tsx agent/simple-rag.ts.

ts
// agent/simple-rag.ts
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

const knowledgeBase: string[] = [
  'This blog is a Next.js app deployed on Vercel. Pushing to the main branch triggers a new production deployment automatically.',
  'Blog posts are stored as JSON files under content/posts and read at build time, so adding a new post requires a redeploy before it appears live.',
  'The blog index page uses Incremental Static Regeneration with a 60 second revalidate window.',
  'Environment variables such as API keys are set in the Vercel dashboard under Project Settings, Environment Variables, never committed to git.',
  'Every pull request gets its own preview deployment with a unique URL, separate from the production domain.',
];

function tokenize(text: string): string[] {
  return text.toLowerCase().match(/[a-z0-9]+/g) ?? [];
}

function vectorize(text: string, vocabulary: string[]): number[] {
  const counts = new Map<string, number>();
  for (const token of tokenize(text)) counts.set(token, (counts.get(token) ?? 0) + 1);
  return vocabulary.map((word) => counts.get(word) ?? 0);
}

function cosineSimilarity(a: number[], b: number[]): number {
  let dot = 0;
  let normA = 0;
  let normB = 0;
  for (let i = 0; i < a.length; i++) {
    dot += a[i] * b[i];
    normA += a[i] * a[i];
    normB += b[i] * b[i];
  }
  if (normA === 0 || normB === 0) return 0;
  return dot / (Math.sqrt(normA) * Math.sqrt(normB));
}

function retrieve(query: string, topK = 2, minScore = 0.22): string[] {
  const vocabulary = Array.from(new Set([query, ...knowledgeBase].flatMap(tokenize)));
  const queryVector = vectorize(query, vocabulary);
  return knowledgeBase
    .map((doc) => ({ doc, score: cosineSimilarity(queryVector, vectorize(doc, vocabulary)) }))
    .sort((a, b) => b.score - a.score)
    .filter((s) => s.score >= minScore)
    .slice(0, topK)
    .map((s) => s.doc);
}

async function askWithContext(question: string): Promise<string> {
  const context = retrieve(question);

  if (context.length === 0) {
    return "I don't have anything in the knowledge base about that.";
  }

  const response = await client.messages.create({
    model: 'claude-sonnet-5',
    max_tokens: 512,
    system:
      'Answer only using the context below. If the context does not contain the answer, say you do not know.\n\nContext:\n' +
      context.map((c, i) => `${i + 1}. ${c}`).join('\n'),
    messages: [{ role: 'user', content: question }],
  });

  const text = response.content.find((b): b is Anthropic.TextBlock => b.type === 'text');
  return text?.text ?? '';
}

askWithContext('How do new blog posts actually go live on the site?').then(console.log);

tokenize and vectorize stand in for a real embedding model. cosineSimilarity and retrieve are the retrieval step exactly as a production system would do it, just against five strings instead of five million. askWithContext is the augmentation step: it builds a system prompt out of whatever retrieve found and only then calls the model.

The Mistake That Took Some Tuning to Find

The first version of retrieve had no minScore at all — it just took the top two matches, whatever their score. I tested it with an unrelated question, something about cooking, that has nothing to do with a blog's deployment setup. It still returned two chunks. The similarity score wasn't zero, because a handful of common words like “is” and “the” overlapped between the question and the documents, and cosine similarity doesn't know those words carry no meaning. The model then answered the cooking question using deployment notes as if they were relevant context, calmly and confidently.

The fix wasn't complicated — reject anything below a minimum score — but picking that number took actual testing, not a guess. I ran the four questions I cared about, printed every score, and looked for a line I could draw between “genuinely related” and “just sharing a few common words.” For this tiny knowledge base, 0.22 sat right between them: every on-topic question still gets its right answer, and the cooking question now correctly returns nothing instead of two irrelevant chunks. That number isn't universal — it's tuned to this exact data, and you'd retune it for a different knowledge base and a different embedding method.

Common Mistakes With RAG

  • Trusting a fixed top-k with no score floor. Returning the two best matches when nothing actually matches well hands the model irrelevant context that it will use anyway. Threshold on the score, not just the rank.
  • Chunking too large or too small. A whole page as one chunk buries the relevant sentence in noise; a single sentence per chunk loses the surrounding context it needs to make sense. Chunk by paragraph or section as a starting point, then adjust based on what actually gets retrieved wrong.
  • Never re-checking what got retrieved. Log the retrieved chunks alongside the answer during development. If the model gives a wrong answer, the first question is whether retrieval handed it the wrong text before you touch the prompt at all.
  • Skipping a real embedding model in production. Word-overlap vectors are fine for learning the shape of RAG. They miss synonyms and paraphrasing entirely — “taka” and “BDT” look unrelated to a word-count vector. A real embedding model, and a proper vector store once the document count grows, are what make retrieval actually reliable.

What's Next

Tool calling gives an agent hands, and RAG gives it a way to know things beyond its training. What's still missing is how to turn a single instruction into a plan the agent actually follows step by step — that's worth its own post.

Frequently Asked Questions

Is RAG the same as fine-tuning?

No. Fine-tuning changes the model's weights through additional training, which is slow to update and better suited to teaching a model a new writing style or behavior. RAG leaves the model untouched and fetches relevant text at the moment of the question, which makes it far cheaper to keep current when the underlying facts change often.

Why not just paste my whole document into the prompt instead of doing retrieval?

For a handful of pages, that's often simpler and you don't need RAG at all. Retrieval starts paying off once your knowledge base is too large to fit in the context window, or once most of it is irrelevant to any single question — retrieval sends only the relevant slice instead of the whole thing every time.

Do I need a real vector database to try this?

Not to learn the mechanism. An array and a similarity function, like the example above, is enough to see how retrieval and generation fit together. Move to a real embedding model and a proper vector database once you have more documents than fit comfortably in memory, or once the process needs to survive a server restart.

img-2
img-2

img1
img1

thumbnail
thumbnail

🟢 Available for Freelance & Contract Work

Need a High-Performance Web App or Custom AI Solution?

I help founders, businesses, and engineering teams build lightning-fast web applications, resilient backend architectures, and intelligent AI workflows. Have an idea in mind? Let’s bring it to life.

Fast MVP Launch (2–4 weeks)
AI Agent & LLM API Integration
100/100 Core Web Vitals & SEO
Production-Ready Clean Architecture

Found this article helpful?

Give some claps to support more in-depth engineering logs!

0 views
Z

Zahid Hasan Tonmoy

Author & Developer

MERN Full Stack Developer & AI Agent Developer based in Dhaka, Bangladesh. Writing about web development, React, PostgreSQL and my learning journey.

Related Articles

Stay Updated

Subscribe to get insights on full-stack architecture, AI agent engineering, and web development.