
Last post was about what an agent remembers between turns and sessions. None of that matters much if the prompt telling it what to do and what it has access to is vague — a perfect memory system can't save an agent that doesn't know when to reach for the tool sitting right in front of it.
Prompt engineering for an agent is different from prompting a plain chatbot. Instead of writing instructions for one response, you're writing something the model re-reads on every single step of a loop, covering its role, exactly when to reach for each tool, where its boundaries are, and what counts as being done. Get a tool-use trigger vague and the model guesses; get the stopping condition vague and it either quits too early or keeps going longer than it should.
A Plain Prompt vs an Agent's System Prompt
A plain prompt asks for one thing and gets judged once: “write a product description,” read, done. An agent's system prompt isn't read once. The post on an agent's four parts called the system prompt a job description, and that holds here too — except this job description gets handed to the model fresh on every single step of the loop, for as long as that loop runs. A vague instruction in a one-shot prompt costs you one bad answer. The same vagueness in an agent's system prompt costs you one bad decision per step, and it doesn't average out over a longer conversation — it compounds.
flowchart LR
SP[System prompt text] --> T1[Read fresh on step 1]
SP --> T2[Read fresh on step 2]
SP --> T3[Read fresh on step N]
T1 --> R1{Ambiguous trigger?}
T2 --> R2{Ambiguous trigger?}
T3 --> R3{Ambiguous trigger?}
R1 -->|risk of a wrong guess| X1[Possible mistake]
R2 -->|risk of a wrong guess| X2[Possible mistake]
R3 -->|risk of a wrong guess| X3[Possible mistake]Flow summary: The same system prompt text is read fresh on every step of the loop, so any ambiguity in it carries the same risk of a wrong guess every single time, not just once — across many steps, that risk adds up instead of averaging out.

The Five Things a Good Agent Prompt Covers
Role. What the agent is, in a sentence or two. Not a biography — just enough that the model knows what kind of task it's being asked to do and what's outside that scope.
Tools, and exactly when to use each one. A list of tool names isn't enough; the model needs to know what situation calls for each one. “Use this whenever the user asks X” is the shape that works, and it matters most precisely when two tools could plausibly apply to the same message.
Boundaries. What the agent must never do: invent data a tool didn't actually return, take an irreversible action without confirmation, or operate outside the tools it's been given. This is the same spirit as the step limits and read-before-write ordering from earlier posts, written into the prompt itself instead of left implicit.
Handling uncertainty. Ask, guess, or escalate — pick one explicitly for ambiguous cases, rather than leaving the model to decide on its own every time it's unsure.
What “done” looks like. An explicit stopping condition, so the agent knows when to stop calling tools and just answer. Without this, you get the same problem a missing step limit causes: the agent either stops before finishing or keeps going past the point of being useful.
Tool Descriptions Are Prompts Too
The tool-calling post already made this point once: the model never reads a tool's implementation, only its description, so that field is itself a piece of prompt engineering, not a separate concern. Going back through this series' actual tool descriptions makes the pattern clear. add_expense, from the memory post, reads: “Record one expense. Use it whenever the user says they spent money.” That trigger phrase is doing real work. get_total, from the same file, reads: “Return the exact total of all recorded expenses in BDT.” No trigger phrase at all — and it gets away with it only because nothing else in that small toolset could be confused with “total.”
That's the real rule, not “always add a trigger phrase” as a blanket mandate: the less distinguishable a tool's name is from its neighbors, the more the description has to do the disambiguating work. Rewritten with a second, similar-sounding tool in mind, get_total becomes: “Return the exact sum of every expense recorded so far, in BDT. Use this whenever the user asks for a total, a sum, or how much they've spent overall — not for a single expense or an average.” Nothing about the tool's behavior changed. Only the sentence a future reader — human or model — needs to disambiguate it from something like get_average_expense did.
A Structured System Prompt, and a Way to Check It
Install the SDK with npm install @anthropic-ai/sdk, set ANTHROPIC_API_KEY if you want to wire this into a real agent, save this as agent/prompt-template.ts and run npx tsx agent/prompt-template.ts to see the checker output on its own.
// agent/prompt-template.ts
const SYSTEM_PROMPT = `
You are an expense-tracking assistant for a single user.
## Role
Record expenses the user mentions and answer questions about what they've spent. You do not give financial advice.
## Tools and when to use them
- add_expense: use this whenever the user states they spent money on something.
- get_total: use this whenever the user asks for a total, a sum, or how much they've spent overall — not for a single expense or an average.
- save_note: use this whenever the user states a lasting preference or constraint, not a one-off detail about today.
## Boundaries
- Never invent an expense, a total, or a saved note the tools did not actually return.
- Never delete or change a past expense unless the user explicitly asks to.
## When you're not sure
If a message could mean more than one thing (for example, an amount with no stated purpose), ask one short clarifying question instead of guessing which tool to call.
## When you're done
Once every expense or note the user mentioned has been recorded and any question has been answered, reply in plain text with no further tool calls.
`.trim();
function hasTriggerPhrase(description: string): boolean {
return /\buse (this|it) whenever\b/i.test(description);
}
const toolDescriptions: Record<string, string> = {
add_expense: 'Record one expense. Use it whenever the user says they spent money.',
get_total_before: 'Return the exact total of all recorded expenses in BDT.',
get_total_after:
"Return the exact sum of every expense recorded so far, in BDT. Use this whenever the user asks for a total, a sum, or how much they've spent overall — not for a single expense or an average.",
};
for (const [name, description] of Object.entries(toolDescriptions)) {
console.log(name, '->', hasTriggerPhrase(description) ? 'has trigger phrase' : 'MISSING trigger phrase');
}
console.log('\n--- system prompt preview ---');
console.log(SYSTEM_PROMPT);
SYSTEM_PROMPT is the five sections laid out as actual headings, meant to be passed straight into system: alongside the familiar tool-and-loop pattern from the tool-calling post. hasTriggerPhrase is a small, honest piece of what can actually be checked mechanically: whether a description contains an explicit “use this/it whenever” trigger. Running it against the real get_total description from the memory post and its rewrite shows exactly the gap described above — the original fails the check, the rewrite passes, and nothing about this requires guessing at how a real model would respond to either one.

Common Mistakes in Agent Prompts
- Treating the system prompt as a single block of advice instead of distinct sections. “Be helpful and careful” doesn't tell the model when to call a tool or when to stop. Structure beats volume.
- Writing tool descriptions in isolation from each other. A description that reads fine alone can still collide with a neighboring tool's description once both are in the same list — review them together, not one at a time.
- No explicit stopping condition. Without one, the model has to infer when it's done from context alone, which is exactly the kind of guess a clear prompt is supposed to remove.
- Padding the prompt with generic advice that doesn't change behavior. Every extra sentence gets re-read on every step. If a line wouldn't change what the model does differently, it's not earning its place.
What's Next
A well-structured prompt tells an agent what it's allowed to do and when. Whether it's actually following that prompt well — and how you'd measure that instead of just assuming — is worth a post of its own.
Frequently Asked Questions
Does a well-written prompt replace good tool design?
No. A clear trigger phrase can't fix a tool whose inputs are genuinely ambiguous, and a perfectly designed tool still needs the model told when to reach for it. They solve different halves of the same problem: tool design from the tool-calling post, prompt structure here.
How long should an agent's system prompt be?
As long as it takes to cover the five sections, and no longer. Generic advice like “be helpful” or “think carefully” gets re-read on every step without changing what the model actually does. If a sentence isn't a role statement, a tool trigger, a boundary, an uncertainty rule, or a stopping condition, it's probably not earning its place.
Should I test my system prompt the way I test my code?
The deterministic parts, yes — whether every tool description has a clear trigger is a checkable property, which is exactly what the small checker in this post does. Whether the model actually follows the prompt well in practice is a different kind of testing, closer to evaluation than a unit test, and it's worth treating as its own topic.
Need a High-Performance Web App or Custom AI Solution?
I help founders, businesses, and engineering teams build lightning-fast web applications, resilient backend architectures, and intelligent AI workflows. Have an idea in mind? Let’s bring it to life.
Found this article helpful?
Give some claps to support more in-depth engineering logs!
Zahid Hasan Tonmoy
Author & Developer
MERN Full Stack Developer & AI Agent Developer based in Dhaka, Bangladesh. Writing about web development, React, PostgreSQL and my learning journey.