Skip to main content

API Gateway Pattern Implementation

Design and implement an API gateway with routing, authentication, rate limiting, request transformation, and circuit breaking for a microservices architecture.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are an API infrastructure architect. Help me implement an API gateway for a microservices architecture with 15 backend microservices.

Step 1: Define the gateway responsibilities: request routing to 15 backend services, authentication and authorization (validating JWT Bearer tokens), rate limiting per client and endpoint, request/response transformation, response caching, and circuit breaking for unhealthy services. Choose between Kong, AWS API Gateway, or custom Node.js based on the project requirements and justify the decision.

Step 2: Design the routing configuration. Create a declarative route mapping that matches incoming requests by path prefix, HTTP method, and headers to the appropriate backend service. Implement path rewriting (strip the service prefix before forwarding), header injection (add X-Request-ID, X-Correlation-ID), and query parameter validation. Support both REST and gRPC protocols.

Step 3: Implement the authentication middleware. Validate JWT Bearer tokens on every request, extract user claims (user ID, roles, permissions), and inject them as trusted headers to backend services. Cache token validation results in Redis for 5 minutes to avoid hitting the auth server on every request. Handle token refresh gracefully by returning 401 with a refresh hint header.

Step 4: Build the rate limiting layer using sliding window log algorithm. Configure tiered limits: 60 requests per minute for free tier, 1,000 for paid tier, and custom limits for enterprise clients. Store rate limit counters in Redis with sliding window accuracy. Return standard rate limit headers (X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset) on every response.

Step 5: Implement circuit breaking for each backend service. Monitor error rates over a 60 seconds sliding window. When the error rate exceeds 50%, open the circuit and return cached responses or graceful fallbacks for 30 seconds before attempting a half-open probe. Log all circuit state transitions and alert the operations team.

Step 6: Design the observability layer. Log every request with: client ID, route, backend service, response status, latency, and rate limit status. Emit metrics for gateway throughput, per-service latency percentiles, error rates, cache hit ratios, and circuit breaker states. Export traces to Jaeger with propagated trace context headers.

What this prompt does

This prompt makes the AI an API infrastructure architect implementing an API gateway for a [architecture_type] architecture with [service_count] backend services. It defines gateway responsibilities, routing configuration, authentication middleware, rate limiting, circuit breaking, and observability. The gateway validates [auth_type] tokens, caches validation results in [cache_layer] for [token_cache_ttl], and applies tiered rate limits via the [rate_limit_algorithm] algorithm.

The structure works because a gateway's job is to consolidate cross-cutting concerns so each of the [service_count] microservices doesn't reimplement them. Routing maps requests by path, method, and headers to the right backend with prefix stripping, header injection, and correlation-ID propagation. The circuit breaker monitors error rates over a [window_duration] window, opens above [error_threshold], and stays open for [open_duration] before a half-open probe — which is what stops one slow service from cascading into the rest. Observability ties it together by logging every request with client, route, status, and latency, emitting per-service metrics, and exporting traces to [tracing_system] with propagated trace context.

When to use it

  • Your microservices need routing, auth, rate limiting, and circuit breaking in one place
  • You want [auth_type] validation centralized instead of duplicated per service
  • You need tiered rate limits separating free and paid clients
  • You want circuit breaking so one unhealthy backend can't drag down the rest
  • You need correlation IDs and traces propagated across services
  • You're choosing between gateway technologies and want the tradeoffs reasoned out
  • You need to support an extra protocol like [additional_protocol] alongside REST

Example output

Expect a design-plus-implementation walkthrough: a responsibilities section justifying the choice among [gateway_option], a declarative routing config with path rewriting and X-Request-ID/X-Correlation-ID injection, authentication middleware that validates [auth_type] tokens, extracts claims, and caches results in [cache_layer] for [token_cache_ttl], a rate-limiting layer with tiered limits and standard X-RateLimit-* headers, per-service circuit-breaker logic with logged state transitions, and an observability section emitting metrics and traces to [tracing_system]. It's structured as labeled steps you can adapt to your stack.

Pro tips

  • Pick [gateway_option] against your real constraints; a managed gateway and a custom one have very different operational costs
  • Cache token validation in [cache_layer] for [token_cache_ttl] to avoid hammering the auth server on every request
  • Tune the circuit breaker's [error_threshold] and [open_duration] per service; one size rarely fits a mixed backend
  • Inject correlation IDs at the gateway so a single request is traceable across all [service_count] services
  • Return standard rate-limit headers consistently; clients rely on them to back off gracefully
  • Support [additional_protocol] at the gateway only if backends actually need it, rather than adding it speculatively
  • Don't skip observability — a gateway you can't see into becomes the hardest part of the system to debug

Frequently Asked Questions

How does the circuit breaker prevent cascading failures?
It monitors each backend's error rate over a `[window_duration]` sliding window and opens the circuit when errors exceed `[error_threshold]`, returning cached responses or fallbacks for `[open_duration]` before a half-open probe. This stops the gateway from piling requests onto a struggling service, which is how one slow backend would otherwise take down the rest.
Does validating tokens at the gateway slow down every request?
It adds a check, but the prompt caches validation results in `[cache_layer]` for `[token_cache_ttl]` so repeat requests skip the auth server. The gateway then injects extracted claims as trusted headers to backends, so services don't re-validate, keeping the per-request cost low.
Can I run this as a managed gateway or does it have to be custom?
Either. The `[gateway_option]` variable covers managed products and a custom build, and step one asks the AI to justify the choice against your requirements. The routing, auth, rate-limiting, and circuit-breaking designs apply regardless, though the exact configuration syntax differs by platform.
How are different client tiers rate-limited differently?
The rate-limiting layer uses the `[rate_limit_algorithm]` algorithm with tiered limits — for example a free-tier and paid-tier request budget per minute, plus custom enterprise limits. Counters live in `[cache_layer]`, and every response carries standard `X-RateLimit-*` headers so clients can see their remaining allowance.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in API Development Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support