What this prompt does
This prompt drives an AI assistant to architect and implement a complete agent system from scratch — not a toy example, but something with the nine structural components that actually matter in production: architecture, tool schemas, system prompt, memory, planning, error recovery, safety, evaluation, and working code. The template's explicit numbered structure forces the model to cover each layer systematically rather than defaulting to a shallow implementation that skips error handling or memory.
The reason it works is that it locks in the hard parts upfront via variables: you define the reasoning strategy, memory type, framework, and language before a single line of code is generated. That prevents the common failure mode where the model produces a ReAct loop with no guardrails, no fallbacks, and no cost cap — which looks impressive in a demo and breaks on day two.
The final instruction — "handle edge cases like ambiguous requests, tool failures, and tasks outside scope" — is load-bearing. It directly prompts the model to generate the defensive paths most developers only add after their first production incident.
When to use it
- You are building a research assistant that must call external APIs, persist context across sessions, and gracefully handle rate limits or API downtime.
- You need a customer-support agent with strict scope limits — only answer questions about your product, escalate everything else, never hallucinate a refund policy.
- You are evaluating LangChain vs CrewAI vs the Anthropic SDK for a new project and want a like-for-like implementation to compare against your requirements.
- You want a multi-agent pipeline (planner + executor + critic) and need the handoff logic and shared memory defined correctly from the start.
- You are shipping an internal tool-use agent and need a test case suite that covers ambiguous inputs and tool timeouts before it goes near users.
Example output
For agent_purpose = "code review assistant", framework = "Anthropic SDK", memory_type = "conversation buffer + vector store for past reviews", language = "Python":
# Tool schema — GitHub PR diff fetcher
{
"name": "fetch_pr_diff",
"description": "Fetch the unified diff for a GitHub pull request.",
"input_schema": {
"type": "object",
"properties": {
"repo": {"type": "string", "description": "owner/repo"},
"pr_number": {"type": "integer"}
},
"required": ["repo", "pr_number"]
}
}
# Error recovery: if fetch_pr_diff fails with 403,
# agent falls back to asking user to paste the diff directly.
# Cost cap: max 4 tool calls per review session.
The system prompt section includes lines like: "You are a senior engineer reviewing this PR for correctness and maintainability. Do not suggest architectural rewrites unless the user explicitly asks. Return findings as a structured list with severity: blocker | suggestion | nit." — a specific persona with a guardrail and an output contract baked in, not a vague instruction to 'be helpful.'
The implementation comes out as an async Python class using anthropic.AsyncAnthropic, with a run_review() entry point, typed tool-call dispatch, and a MemoryStore wrapping both an in-memory buffer and a vector index for retrieving context from past reviews of the same repo.
Pro tips
- Set
reasoning_typeexplicitly. "ReAct" and "plan-and-execute" produce very different code. ReAct re-plans after every tool call inline; plan-and-execute generates a full plan first, then executes. For multi-step tasks with uncertain tool availability, plan-and-execute gives the agent a chance to validate its plan before burning API calls. - Be precise with
memory_type. "Short-term" generates an in-memory buffer that evaporates on restart. If you need persistence, write "Redis-backed conversation history + pgvector for semantic recall" — the model will implement both, including the connection and schema. - Name your tools after their failure modes. In
available_tools, add notes like "search_web (may time out after 5s, returns empty list on failure)" — the model uses this to write better retry logic and fallback branches, not just a generic try/except. - Harden the evaluation step by specifying the failure taxonomy. Step 8 generates test cases, but if you add "include tests for: tool timeout, malformed tool response, user request outside agent scope, and contradictory instructions" the coverage jumps from happy-path to something you can actually ship against.
- Pair with a cost audit. After getting the implementation, follow up with: "Show me the estimated token cost per agent run and where the biggest cost drivers are." Agents are expensive to run wrong — the audit often reveals a memory retrieval step that doubles costs for no measurable quality gain.