What this prompt does
This prompt asks the AI to design a full autonomous agent system around a single [agent_purpose], built on [model_provider] and implemented in [language] with [framework]. Instead of returning a toy loop, it walks through the ten parts that actually decide whether an agent survives production: the observe-think-act loop capped at [max_iterations], typed tool schemas for [tools], a three-tier memory system, a planning module for [task_type], error recovery, and guardrails.
The structure works because it forces the model to reason about failure modes before happy paths. By naming [dangerous_actions] to block and [confirm_actions] that require a human in the loop, you push the design toward safety rather than treating it as an afterthought. The cost-tracking, logging, and [eval_count]-scenario evaluation harness steps mean you get an architecture you can debug and trust, not just a demo that works once.
The three-tier memory split is what keeps a long-running agent coherent. Short-term memory holds the live conversation, working memory tracks the current task, and long-term memory in [memory_backend] persists what matters across runs. Pairing that with a planning module that decomposes [task_type] work into sub-tasks means the agent reasons about the whole job rather than reacting one observation at a time, which is the difference between an agent that finishes a multi-step task and one that wanders.
When to use it
- You are building an agent that takes real actions (files, issues, deployments) and need guardrails baked in from the start.
- You want a reference architecture before writing code, so the team agrees on the loop, memory, and tool boundaries.
- You need typed tool schemas and an observation-parsing strategy rather than ad-hoc string handling.
- You are moving an agent from prototype to production and need cost limits, retries, and human escalation.
- You want an evaluation harness so regressions in agent behaviour are caught before users hit them.
Example output
Expect a structured architecture document: an annotated agent-loop diagram, a table of tool definitions with input/output schemas, a description of the short-term, working, and long-term [memory_backend] layers, and a sample run trace showing observe-think-act steps with the guardrails firing. It typically closes with skeleton code in [language] and the evaluation harness layout.
Pro tips
- Make
[agent_purpose]narrow and verb-driven. "Automates code review and creates issues" produces a far cleaner design than a vague "helps with engineering." - Keep
[max_iterations]low at first. I cap it tight so a confused agent fails fast instead of burning budget in a loop. - Be explicit and generous in
[dangerous_actions]and[confirm_actions]— anything irreversible (deletes, pushes to main) belongs there. - List
[tools]as concrete function names with clear verbs; the model writes better schemas when the tool intent is obvious. - Match
[memory_backend]to your real stack. A vector database is overkill if simple structured storage covers your[task_type]. - Treat the
[eval_count]scenarios as a starting point and expand them from real failures once the agent runs. - Wire in cost tracking and a per-run budget early, since an agent that loops without a budget cap can quietly burn money before you notice.
- Keep observation parsing structured rather than regexing raw tool output; the planning module reasons far better over typed data.