What this prompt does
This prompt turns the model into an AI engineer who has shipped LLM features at scale, then walks it through a seven-part architecture review for the feature you describe in [feature_description]. Instead of a vague "use an LLM" answer, it forces concrete decisions: model routing tables, a prompt-management system, caching layers, rate limiting, resilience patterns, monitoring, and starter implementation code in your [tech_stack].
The structure works because every section is anchored to your real constraints. [usage_volume], [latency_requirement], and [monthly_budget] are what separate a toy demo from a system that survives production traffic. By feeding the model your numbers, you get a routing table that splits simple tasks to cheaper models and complex tasks to larger ones, a cache-hit estimate tied to your actual feature, and a [cost_concern] section that addresses the specific failure mode you're worried about rather than generic advice.
When to use it
- You're about to commit to an LLM-feature design and want to pressure-test it before writing code.
- Your prototype works but you have no caching, routing, or cost controls and traffic is coming.
- You need to justify model and infrastructure choices to a team or client with real cost-per-token math.
- A feature is blowing past its
[monthly_budget]and you need a routing and caching strategy to rein it in. - You want starter implementation code (client wrapper, prompt manager, response parser) in your stack to copy from.
- You're handling unpredictable input sizes and need a token-limit strategy spelled out in
[cost_concern].
Example output
You get a structured architecture document: a model-selection table mapping each task to a recommended model with cost and latency columns, a prompt-versioning scheme, exact-match plus semantic caching design with a cache-hit estimate, rate-limit and token-budget rules, a resilience section (retries, circuit breaker, graceful degradation), a monitoring plan with p50/p95/p99 and quality metrics, and code scaffolding in your [tech_stack] — an LLM client wrapper, prompt template manager, response parser, and cost-tracking middleware with mocked tests.
Pro tips
- Be specific in
[feature_description]— "code review assistant for PR diffs" yields a far better routing table than "AI helper." - Put real figures in
[usage_volume]and[monthly_budget]; the caching and cost sections are only as good as the numbers you feed them. - Use
[cost_concern]to name your scariest edge case (huge inputs, abuse, spiky traffic) so the model designs around it. - Treat the generated code as scaffolding, not finished work — validate the retry, timeout, and cache-invalidation logic against your own infrastructure.
- Re-run with a tighter
[latency_requirement]to see how the routing and streaming recommendations shift. - Ask a follow-up to expand any one section (semantic caching, the evaluation pipeline) into full implementation once the high-level design looks right.