Skip to main content

Claude Prompt to Build a Production AI Agent Workflow

Build production AI agents with tool use, memory, planning, and safety guardrails using the Anthropic SDK, LangChain, or CrewAI, with complete code.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
Design an AI agent system for automated code review and PR feedback bot.

User interaction: GitHub webhook triggers + comment responses
Available tools/APIs: GitHub API, code analysis, file reading, web search
Framework: Anthropic Agent SDK

Build the complete agent:
1. **Agent Architecture** — single agent vs multi-agent, reasoning strategy (ReAct (Reasoning + Acting))
2. **Tool Definitions** — schema for each tool with input validation and error handling
3. **System Prompt** — detailed instructions, personality, guardrails, and output format
4. **Memory System** — conversation buffer with summary compression for maintaining context across conversations
5. **Planning Strategy** — how the agent breaks down complex tasks
6. **Error Recovery** — retry logic, fallbacks, graceful degradation
7. **Safety Guardrails** — input validation, output filtering, rate limiting, cost caps
8. **Evaluation** — test cases to verify agent behavior
9. **Complete Code** — working Python implementation

The agent should handle edge cases like: ambiguous user requests, tool failures, and tasks outside its scope. Make it production-ready, not a demo.

What this prompt does

This prompt drives an AI assistant to architect and implement a complete agent system from scratch — not a toy example, but something with the nine structural components that actually matter in production: architecture, tool schemas, system prompt, memory, planning, error recovery, safety, evaluation, and working code. The template's explicit numbered structure forces the model to cover each layer systematically rather than defaulting to a shallow implementation that skips error handling or memory.

The reason it works is that it locks in the hard parts upfront via variables: you define the reasoning strategy, memory type, framework, and language before a single line of code is generated. That prevents the common failure mode where the model produces a ReAct loop with no guardrails, no fallbacks, and no cost cap — which looks impressive in a demo and breaks on day two.

The final instruction — "handle edge cases like ambiguous requests, tool failures, and tasks outside scope" — is load-bearing. It directly prompts the model to generate the defensive paths most developers only add after their first production incident.

When to use it

  • You are building a research assistant that must call external APIs, persist context across sessions, and gracefully handle rate limits or API downtime.
  • You need a customer-support agent with strict scope limits — only answer questions about your product, escalate everything else, never hallucinate a refund policy.
  • You are evaluating LangChain vs CrewAI vs the Anthropic SDK for a new project and want a like-for-like implementation to compare against your requirements.
  • You want a multi-agent pipeline (planner + executor + critic) and need the handoff logic and shared memory defined correctly from the start.
  • You are shipping an internal tool-use agent and need a test case suite that covers ambiguous inputs and tool timeouts before it goes near users.

Example output

For agent_purpose = "code review assistant", framework = "Anthropic SDK", memory_type = "conversation buffer + vector store for past reviews", language = "Python":

# Tool schema — GitHub PR diff fetcher
{
  "name": "fetch_pr_diff",
  "description": "Fetch the unified diff for a GitHub pull request.",
  "input_schema": {
    "type": "object",
    "properties": {
      "repo": {"type": "string", "description": "owner/repo"},
      "pr_number": {"type": "integer"}
    },
    "required": ["repo", "pr_number"]
  }
}

# Error recovery: if fetch_pr_diff fails with 403,
# agent falls back to asking user to paste the diff directly.
# Cost cap: max 4 tool calls per review session.

The system prompt section includes lines like: "You are a senior engineer reviewing this PR for correctness and maintainability. Do not suggest architectural rewrites unless the user explicitly asks. Return findings as a structured list with severity: blocker | suggestion | nit." — a specific persona with a guardrail and an output contract baked in, not a vague instruction to 'be helpful.'

The implementation comes out as an async Python class using anthropic.AsyncAnthropic, with a run_review() entry point, typed tool-call dispatch, and a MemoryStore wrapping both an in-memory buffer and a vector index for retrieving context from past reviews of the same repo.

Pro tips

  • Set reasoning_type explicitly. "ReAct" and "plan-and-execute" produce very different code. ReAct re-plans after every tool call inline; plan-and-execute generates a full plan first, then executes. For multi-step tasks with uncertain tool availability, plan-and-execute gives the agent a chance to validate its plan before burning API calls.
  • Be precise with memory_type. "Short-term" generates an in-memory buffer that evaporates on restart. If you need persistence, write "Redis-backed conversation history + pgvector for semantic recall" — the model will implement both, including the connection and schema.
  • Name your tools after their failure modes. In available_tools, add notes like "search_web (may time out after 5s, returns empty list on failure)" — the model uses this to write better retry logic and fallback branches, not just a generic try/except.
  • Harden the evaluation step by specifying the failure taxonomy. Step 8 generates test cases, but if you add "include tests for: tool timeout, malformed tool response, user request outside agent scope, and contradictory instructions" the coverage jumps from happy-path to something you can actually ship against.
  • Pair with a cost audit. After getting the implementation, follow up with: "Show me the estimated token cost per agent run and where the biggest cost drivers are." Agents are expensive to run wrong — the audit often reveals a memory retrieval step that doubles costs for no measurable quality gain.

Frequently Asked Questions

Can I use this prompt for a multi-agent setup, not just a single agent?
Yes — step 1 of the template explicitly asks the model to decide between single-agent and multi-agent architecture. If you set agent_purpose to something like 'market research pipeline with a planner, web researcher, and summarizer,' the model will generate separate agent definitions with handoff logic. Specify your preferred orchestration pattern (sequential, hierarchical, or parallel) in the interaction_mode or reasoning_type field to get the topology you want rather than letting the model default to sequential.
Which framework should I pick — LangChain, CrewAI, or the Anthropic SDK?
It depends on what you are optimizing for. LangChain gives you the widest tool ecosystem and the most existing tutorials, but adds abstraction overhead that can obscure what is actually happening in tool calls. CrewAI is purpose-built for multi-agent role assignment and is the fastest path if you have clearly defined agent roles. The Anthropic SDK gives you the most direct control over tool schemas and is the right choice if you want minimal dependencies, predictable token costs, and access to the latest model capabilities without waiting for a framework update. Run this prompt once per framework on the same agent_purpose — the side-by-side comparison will reveal tradeoffs no benchmark captures.
Will this generate code that is actually safe to run in production, or do I need to add safety layers myself?
The template explicitly asks for safety guardrails (step 7: input validation, output filtering, rate limiting, cost caps) and error recovery (step 6), so the generated code will include these layers. That said, treat the output as a strong starting point, not a finished product. Review the generated cost cap logic — models sometimes hard-code a token limit that is too low for real workloads. Also verify that output filtering matches your actual content policy, since the model will infer guardrails from context rather than knowing your specific requirements. Run the generated test cases from step 8 against your actual tool endpoints before deploying.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in AI & Machine Learning Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support