Skip to main content

Claude Prompt to Build an AI Agent with Tool Use & Memory

Design an autonomous AI agent with tool use, memory, planning, error recovery, and guardrails. Production-ready architecture in any language.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
Design an AI agent system that automates code review and creates GitHub issues for findings using Claude API with tool_use. Architecture: 1) Agent loop: observe → think → act → observe with max 10 iterations, 2) Tool definitions for read_file, search_code, create_issue, run_tests, get_diff with typed input/output schemas, 3) Memory system — short-term (conversation), working (current task), long-term (vector database (Pinecone)), 4) Planning module that breaks multi-file code review with dependency analysis tasks into sub-tasks, 5) Error recovery — retry strategies, fallback tools, and human escalation, 6) Guardrails — prevent deleting files, pushing to main, modifying CI/CD, require confirmation for creating issues, commenting on PRs, 7) Observation parsing — extract structured information from tool outputs, 8) Cost tracking and budget limits per agent run, 9) Logging and tracing for debugging agent decisions, 10) Evaluation harness with 10 test scenarios. Implement in TypeScript using Anthropic Claude Agent SDK. Include architecture diagram and example agent run trace.

What this prompt does

This prompt asks the AI to design a full autonomous agent system around a single [agent_purpose], built on [model_provider] and implemented in [language] with [framework]. Instead of returning a toy loop, it walks through the ten parts that actually decide whether an agent survives production: the observe-think-act loop capped at [max_iterations], typed tool schemas for [tools], a three-tier memory system, a planning module for [task_type], error recovery, and guardrails.

The structure works because it forces the model to reason about failure modes before happy paths. By naming [dangerous_actions] to block and [confirm_actions] that require a human in the loop, you push the design toward safety rather than treating it as an afterthought. The cost-tracking, logging, and [eval_count]-scenario evaluation harness steps mean you get an architecture you can debug and trust, not just a demo that works once.

The three-tier memory split is what keeps a long-running agent coherent. Short-term memory holds the live conversation, working memory tracks the current task, and long-term memory in [memory_backend] persists what matters across runs. Pairing that with a planning module that decomposes [task_type] work into sub-tasks means the agent reasons about the whole job rather than reacting one observation at a time, which is the difference between an agent that finishes a multi-step task and one that wanders.

When to use it

  • You are building an agent that takes real actions (files, issues, deployments) and need guardrails baked in from the start.
  • You want a reference architecture before writing code, so the team agrees on the loop, memory, and tool boundaries.
  • You need typed tool schemas and an observation-parsing strategy rather than ad-hoc string handling.
  • You are moving an agent from prototype to production and need cost limits, retries, and human escalation.
  • You want an evaluation harness so regressions in agent behaviour are caught before users hit them.

Example output

Expect a structured architecture document: an annotated agent-loop diagram, a table of tool definitions with input/output schemas, a description of the short-term, working, and long-term [memory_backend] layers, and a sample run trace showing observe-think-act steps with the guardrails firing. It typically closes with skeleton code in [language] and the evaluation harness layout.

Pro tips

  • Make [agent_purpose] narrow and verb-driven. "Automates code review and creates issues" produces a far cleaner design than a vague "helps with engineering."
  • Keep [max_iterations] low at first. I cap it tight so a confused agent fails fast instead of burning budget in a loop.
  • Be explicit and generous in [dangerous_actions] and [confirm_actions] — anything irreversible (deletes, pushes to main) belongs there.
  • List [tools] as concrete function names with clear verbs; the model writes better schemas when the tool intent is obvious.
  • Match [memory_backend] to your real stack. A vector database is overkill if simple structured storage covers your [task_type].
  • Treat the [eval_count] scenarios as a starting point and expand them from real failures once the agent runs.
  • Wire in cost tracking and a per-run budget early, since an agent that loops without a budget cap can quietly burn money before you notice.
  • Keep observation parsing structured rather than regexing raw tool output; the planning module reasons far better over typed data.

Frequently Asked Questions

Does this prompt actually write the agent code or just the architecture?
It produces both an architecture design and skeleton implementation code in your chosen `[language]` and `[framework]`. The depth of the code depends on the model and how specific your variables are, but expect a working scaffold rather than a fully tested system you can ship untouched.
Can I use a provider other than Claude for the agent?
Yes. The `[model_provider]` variable is open, so you can target any API that supports tool use. The loop, memory, and guardrail patterns are provider-agnostic, though tool-calling syntax and schema formats differ between providers and will need adjusting.
Why does the prompt insist on guardrails and confirmation steps?
Because an agent that can delete files or push to main without checks is dangerous in production. Separating `[dangerous_actions]` to block from `[confirm_actions]` that need human approval is what makes an autonomous agent safe to actually run against real systems.
What is the evaluation harness for?
It defines `[eval_count]` test scenarios so you can verify agent behaviour repeatably. This catches regressions when you change prompts, tools, or models, which is essential because agent behaviour is non-deterministic and silent quality drops are easy to miss otherwise.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in AI Chatbot & Agent Building Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support