Skip to main content

API Rate Limiting with Redis

Implement production-grade API rate limiting using Redis with sliding windows, tiered limits, distributed coordination, and graceful degradation.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a backend engineer specializing in API reliability. Implement a comprehensive rate limiting system using Redis for a Express.js API serving 10,000 requests per second.

Step 1: Implement the sliding window counter rate limiting algorithm in Redis. Write the Lua script that atomically checks and increments the counter for a given key (composed of client identifier + endpoint). The script should return the current count, remaining allowance, and reset timestamp. Use rl:{client_id}:{endpoint}:{window} as the Redis key format. Ensure the TTL is set correctly to auto-expire windows.

Step 2: Design a tiered rate limiting configuration. Define 4 tiers with different limits: anonymous (by IP, 30/min), authenticated free (100/min), authenticated premium (1,000/min), and internal services (no limit but tracked). Store tier configurations in database table and cache them in application memory with a 60 seconds refresh interval.

Step 3: Implement endpoint-specific rate limits that override the global tier limits. For example, POST endpoints get 30 requests per minute, file upload endpoints get 10, and read-only GET endpoints get the full tier limit. Create a configuration DSL that maps route patterns to specific limits.

Step 4: Build the rate limit response handling. On every response, include standard headers: X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset (Unix timestamp), and Retry-After (when limited). When a client exceeds the limit, return HTTP 429 with a JSON body containing the error message, retry_after seconds, and a link to the rate limit documentation.

Step 5: Implement graceful degradation for when Redis is unavailable. Fall back to an in-memory token bucket rate limiter that applies conservative default limits. Log a warning and alert the operations team. When Redis recovers, seamlessly switch back without resetting client counters (use the higher of the two counts).

Step 6: Create a rate limit monitoring dashboard that tracks: requests allowed vs rejected per tier, top clients by request volume, endpoints hitting limits most frequently, Redis latency for rate limit checks, and fallback activation events. Expose these metrics via a Prometheus endpoint for Grafana to scrape.

What this prompt does

This prompt makes the AI a backend engineer implementing a rate limiting system using Redis for a [framework] API serving [rps] requests per second. It covers the core Redis algorithm, tiered configuration, endpoint-specific overrides, response handling, graceful degradation, and monitoring. Step one writes the atomic Lua script for the [algorithm] algorithm using [key_format] keys, returning the current count, remaining allowance, and reset timestamp.

The structure works because production rate limiting has two non-negotiable parts: atomicity and a fallback. The Lua script runs check-and-increment atomically inside Redis so concurrent requests at [rps] scale can't race past the limit, and it returns the current count, remaining allowance, and reset timestamp in one round-trip. The graceful-degradation step adds an in-memory [fallback_algorithm] limiter with conservative defaults for when Redis is unavailable, and when Redis recovers it reconciles by taking the higher of the two counts so clients aren't handed a fresh allowance mid-outage. Tiered limits separate anonymous, free, and premium clients, with per-endpoint overrides like tighter write and upload caps layered on top through a configuration DSL.

When to use it

  • Your API needs rate limiting that holds under real concurrent load
  • You want an atomic Redis implementation rather than a race-prone read-then-write
  • You need tiered limits across anonymous, free, and premium clients
  • You want stricter caps on write and upload endpoints than reads
  • You need a sane fallback for when Redis blips
  • You want standard rate-limit headers and a monitoring view of allowed-versus-rejected traffic
  • You need tier configs in [config_store] cached and refreshed without a lookup per request

Example output

Expect implementation-focused output: the atomic Lua script for [algorithm] with [key_format] keys and correctly set TTLs that auto-expire windows, a tiered configuration loaded from [config_store] and cached in memory with a [config_refresh] interval, an endpoint-override DSL mapping route patterns to specific limits, response handling that sets X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset, and Retry-After headers and returns HTTP 429 with a JSON body, the in-memory [fallback_algorithm] degradation path, and a monitoring section exposing allowed-versus-rejected and Redis-latency metrics in [metrics_format]. It's structured as labeled steps for [framework].

Pro tips

  • Keep the check-and-increment in a single Lua script; splitting it across round-trips reintroduces the race you're trying to avoid
  • Set TTLs correctly so windows auto-expire — a missing TTL silently leaks keys and skews counts
  • Use the [key_format] default's client-plus-endpoint shape so limits are scoped where you expect
  • On Redis recovery, take the higher of the in-memory and Redis counts so clients don't get a free reset mid-outage
  • Cache [config_store] tiers in memory with a [config_refresh] interval to avoid a lookup on every request
  • Always return Retry-After on a 429; well-behaved clients use it to back off instead of retry-storming

Frequently Asked Questions

Why does the rate limiter need a Lua script instead of plain Redis commands?
A Lua script runs the check-and-increment atomically inside Redis, so two concurrent requests can't both read an under-limit count and then both increment past it. Doing the same with separate commands opens a race window that lets clients exceed the limit under `[rps]`-level concurrency.
What happens to rate limiting if Redis goes down?
The graceful-degradation step falls back to an in-memory `[fallback_algorithm]` limiter with conservative defaults and alerts operations. When Redis recovers, the system switches back and reconciles by using the higher of the in-memory and Redis counts, so clients don't get a fresh allowance during the outage.
Can different endpoints have different limits from the tier limit?
Yes. Beyond the tiered limits for anonymous, free, and premium clients, the prompt adds endpoint-specific overrides through a configuration DSL — for instance tighter caps on write and file-upload endpoints than on read-only GETs. Route patterns map to those specific limits.
Which response headers does it return when a client is limited?
Every response includes `X-RateLimit-Limit`, `X-RateLimit-Remaining`, and `X-RateLimit-Reset`, and a limited request also returns `Retry-After` with an HTTP 429 and a JSON body. The `Retry-After` value tells well-behaved clients exactly how long to wait before retrying.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in API Development Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support