What this prompt does
This prompt makes the AI an API infrastructure architect implementing an API gateway for a [architecture_type] architecture with [service_count] backend services. It defines gateway responsibilities, routing configuration, authentication middleware, rate limiting, circuit breaking, and observability. The gateway validates [auth_type] tokens, caches validation results in [cache_layer] for [token_cache_ttl], and applies tiered rate limits via the [rate_limit_algorithm] algorithm.
The structure works because a gateway's job is to consolidate cross-cutting concerns so each of the [service_count] microservices doesn't reimplement them. Routing maps requests by path, method, and headers to the right backend with prefix stripping, header injection, and correlation-ID propagation. The circuit breaker monitors error rates over a [window_duration] window, opens above [error_threshold], and stays open for [open_duration] before a half-open probe — which is what stops one slow service from cascading into the rest. Observability ties it together by logging every request with client, route, status, and latency, emitting per-service metrics, and exporting traces to [tracing_system] with propagated trace context.
When to use it
- Your microservices need routing, auth, rate limiting, and circuit breaking in one place
- You want
[auth_type]validation centralized instead of duplicated per service - You need tiered rate limits separating free and paid clients
- You want circuit breaking so one unhealthy backend can't drag down the rest
- You need correlation IDs and traces propagated across services
- You're choosing between gateway technologies and want the tradeoffs reasoned out
- You need to support an extra protocol like
[additional_protocol]alongside REST
Example output
Expect a design-plus-implementation walkthrough: a responsibilities section justifying the choice among [gateway_option], a declarative routing config with path rewriting and X-Request-ID/X-Correlation-ID injection, authentication middleware that validates [auth_type] tokens, extracts claims, and caches results in [cache_layer] for [token_cache_ttl], a rate-limiting layer with tiered limits and standard X-RateLimit-* headers, per-service circuit-breaker logic with logged state transitions, and an observability section emitting metrics and traces to [tracing_system]. It's structured as labeled steps you can adapt to your stack.
Pro tips
- Pick
[gateway_option]against your real constraints; a managed gateway and a custom one have very different operational costs - Cache token validation in
[cache_layer]for[token_cache_ttl]to avoid hammering the auth server on every request - Tune the circuit breaker's
[error_threshold]and[open_duration]per service; one size rarely fits a mixed backend - Inject correlation IDs at the gateway so a single request is traceable across all
[service_count]services - Return standard rate-limit headers consistently; clients rely on them to back off gracefully
- Support
[additional_protocol]at the gateway only if backends actually need it, rather than adding it speculatively - Don't skip observability — a gateway you can't see into becomes the hardest part of the system to debug