Skip to main content

Claude Prompt to Design a Production LLM Application Architecture

Design production LLM apps with model routing, prompt versioning, caching, rate limits, resilience, monitoring, and evaluation, with implementation code.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are an AI engineer who has deployed LLM applications at scale. Help me design a production-grade architecture for my LLM-powered feature.

**Application Context:**
- Feature: AI-powered code review assistant that analyzes PR diffs and provides actionable feedback on code quality, security, and performance
- Expected usage: 500 code reviews per day, average 2K tokens input, 1K tokens output per review
- Latency requirement: Under 10 seconds for initial feedback, streaming preferred
- Budget: $500/month for LLM API costs
- Tech stack: Python FastAPI backend, Redis for caching, PostgreSQL for persistence

**Design These Architecture Components:**

**1. LLM Selection & Routing:**
Based on AI-powered code review assistant that analyzes PR diffs and provides actionable feedback on code quality, security, and performance, recommend:

| Task | Recommended Model | Why | Cost per 1K tokens | Latency |
|------|------------------|-----|--------------------| ---------|
| [Task 1] | [Model] | [Reason] | [Cost] | [Latency] |
| [Task 2] | [Model] | [Reason] | [Cost] | [Latency] |

**Model routing strategy:**
- Simple tasks → smaller/cheaper model (GPT-4o-mini, Claude Haiku, Llama)
- Complex tasks → larger model (GPT-4o, Claude Sonnet, etc.)
- How to classify requests to route them correctly
- Fallback chain: if primary model is down, route to [Backup Model]

**2. Prompt Management System:**
- Version-controlled prompt templates (not hardcoded in application code)
- A/B testing framework for prompt variants
- Prompt template structure with [Variables]:
  ```
  System: [versioned system prompt]
  Context: [dynamically injected context]
  User: [user input with guardrails]
  ```
- Input validation and sanitization (prevent prompt injection)
- Output parsing and validation (structured output with retry on malformed responses)

**3. Caching Strategy:**
- **Exact match cache:** Same input → cached output (Redis with TTL)
- **Semantic cache:** Similar inputs → cached output (embedding similarity threshold)
- **Cache invalidation:** When prompts change, invalidate affected cache entries
- Estimated cache hit rate for AI-powered code review assistant that analyzes PR diffs and provides actionable feedback on code quality, security, and performance: [Estimate]%
- Monthly cost savings from caching: [Estimate]

**4. Rate Limiting & Cost Control:**
- Per-user rate limits (prevent abuse)
- Global rate limits (stay within API budget)
- Token budget per request (max input + max output tokens)
- Cost tracking per feature, per user, per model
- Alert when daily/weekly spend exceeds threshold
- Users can submit very large PRs — need to handle token limits gracefully by chunking

**5. Error Handling & Resilience:**
- Timeout handling (LLM calls can be slow)
- Retry strategy with exponential backoff
- Circuit breaker for API outages
- Graceful degradation (what to show users when LLM is unavailable)
- Content filtering (detect and handle inappropriate outputs)

**6. Monitoring & Evaluation:**
- **Operational metrics:** Latency p50/p95/p99, error rate, token usage, cost per request
- **Quality metrics:** User satisfaction (thumbs up/down), task completion rate
- **Drift detection:** Monitor output quality over time (model updates can change behavior)
- Evaluation pipeline: automated tests that run against model outputs
- A/B test framework for comparing models/prompts

**7. Implementation Code:**
Generate the core architecture in Python FastAPI backend, Redis for caching, PostgreSQL for persistence:
- LLM client wrapper with retry, timeout, caching
- Prompt template manager
- Response parser with structured output validation
- Cost tracking middleware
- Integration tests with mocked LLM responses

What this prompt does

This prompt turns the model into an AI engineer who has shipped LLM features at scale, then walks it through a seven-part architecture review for the feature you describe in [feature_description]. Instead of a vague "use an LLM" answer, it forces concrete decisions: model routing tables, a prompt-management system, caching layers, rate limiting, resilience patterns, monitoring, and starter implementation code in your [tech_stack].

The structure works because every section is anchored to your real constraints. [usage_volume], [latency_requirement], and [monthly_budget] are what separate a toy demo from a system that survives production traffic. By feeding the model your numbers, you get a routing table that splits simple tasks to cheaper models and complex tasks to larger ones, a cache-hit estimate tied to your actual feature, and a [cost_concern] section that addresses the specific failure mode you're worried about rather than generic advice.

When to use it

  • You're about to commit to an LLM-feature design and want to pressure-test it before writing code.
  • Your prototype works but you have no caching, routing, or cost controls and traffic is coming.
  • You need to justify model and infrastructure choices to a team or client with real cost-per-token math.
  • A feature is blowing past its [monthly_budget] and you need a routing and caching strategy to rein it in.
  • You want starter implementation code (client wrapper, prompt manager, response parser) in your stack to copy from.
  • You're handling unpredictable input sizes and need a token-limit strategy spelled out in [cost_concern].

Example output

You get a structured architecture document: a model-selection table mapping each task to a recommended model with cost and latency columns, a prompt-versioning scheme, exact-match plus semantic caching design with a cache-hit estimate, rate-limit and token-budget rules, a resilience section (retries, circuit breaker, graceful degradation), a monitoring plan with p50/p95/p99 and quality metrics, and code scaffolding in your [tech_stack] — an LLM client wrapper, prompt template manager, response parser, and cost-tracking middleware with mocked tests.

Pro tips

  • Be specific in [feature_description] — "code review assistant for PR diffs" yields a far better routing table than "AI helper."
  • Put real figures in [usage_volume] and [monthly_budget]; the caching and cost sections are only as good as the numbers you feed them.
  • Use [cost_concern] to name your scariest edge case (huge inputs, abuse, spiky traffic) so the model designs around it.
  • Treat the generated code as scaffolding, not finished work — validate the retry, timeout, and cache-invalidation logic against your own infrastructure.
  • Re-run with a tighter [latency_requirement] to see how the routing and streaming recommendations shift.
  • Ask a follow-up to expand any one section (semantic caching, the evaluation pipeline) into full implementation once the high-level design looks right.

Frequently Asked Questions

Does this prompt write actual code or just describe the architecture?
It does both. The first six sections design the architecture conceptually, and section seven generates starter code in your `[tech_stack]` — a client wrapper, prompt manager, response parser, and cost middleware. Treat that code as scaffolding to validate, not production-ready output.
Will it recommend specific models like GPT-4o or Claude?
Yes. The routing section builds a table that maps each task to a recommended model with cost-per-token and latency columns, typically splitting cheap models for simple tasks and larger ones for complex work. Verify the current pricing yourself, since model costs change frequently.
Can I use it if I haven't picked a tech stack yet?
You can, but the implementation code is weaker without a concrete `[tech_stack]`. Leave it generic for the design sections, then re-run once you've chosen a stack to get usable client and caching code tailored to your framework.
How accurate are the cache-hit-rate and cost-savings estimates?
They are reasoned estimates based on the `[feature_description]` and `[usage_volume]` you provide, not measurements. Use them to compare design options and set expectations, but always confirm real cache performance with production telemetry before trusting the numbers.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in AI & Machine Learning Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support