Skip to main content

Claude/ChatGPT Prompt to Design Real-Time Chat for 10M Concurrent Users

Design a WhatsApp-style real-time chat system handling 10M concurrent connections with delivery guarantees and global latency budgets.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a senior distributed systems architect. Design a real-time chat system end to end and specify it tightly enough to build — concrete components, protocols, and numbers, not hand-waving.

Context:
- Concurrent users: 10M concurrent users
- Connection protocol: WebSocket (fallback to SSE/long-poll)
- Message store: Cassandra for messages, Elasticsearch for search
- Region footprint: us-east, eu-west, ap-south (multi-region)

Deliverables:
1. Connection layer: gateway sharding, sticky routing, and connection-to-node mapping for 10M concurrent users live sockets.
2. Message delivery guarantees (at-least-once vs exactly-once), ordering, and dedup keys.
3. Group chat fan-out strategy (write fan-out vs read fan-out) with the crossover threshold.
4. Storage and search using Cassandra for messages, Elasticsearch for search, plus offline message queueing and read receipts.
5. Presence, E2E encryption hooks, and the global latency budget across us-east, eu-west, ap-south (multi-region).
6. Failure modes: gateway loss, store partition, and backpressure handling.

Return an architecture diagram description, a per-component table, and back-of-envelope capacity math.

What this prompt does

This prompt makes the AI design a real-time chat system end to end, specified tightly enough to build with concrete components, protocols, and numbers. You set the [concurrent_users], the [connection_protocol], the [message_store], and the [region_footprint]. It returns a connection layer with gateway sharding and sticky routing, message delivery guarantees with ordering and dedup keys, a group fan-out strategy with a crossover threshold, storage and search plus offline queueing and read receipts, presence and E2E encryption hooks with a latency budget, and failure modes.

The structure works because the part that bites in real-time systems is fan-out: naive write fan-out to large groups melts your queue. Forcing the connection sharding and store decisions before code pressure-tests the design. [concurrent_users] drives the gateway sharding and connection-to-node mapping math, [connection_protocol] sets the transport, [message_store] shapes storage and search, and [region_footprint] defines the global latency budget.

When to use it

  • You're designing a chat or messaging feature and need the connection layer right.
  • You're prepping a system design interview on real-time systems at scale.
  • You need to reason about gateway sharding and sticky routing for many sockets.
  • You want delivery guarantees, ordering, and dedup keys defined explicitly.
  • You need the group fan-out crossover threshold (write versus read) pinned down.
  • You want presence, offline queueing, read receipts, and named failure modes.

Example output

You get an architecture diagram description, a per-component table, and back-of-envelope capacity math: a connection layer with gateway sharding, sticky routing, and connection-to-node mapping for [concurrent_users] live sockets; message delivery guarantees (at-least-once versus exactly-once) with ordering and dedup keys; a group fan-out strategy (write versus read) with the crossover threshold; storage and search using [message_store] plus offline message queueing and read receipts; presence, E2E encryption hooks, and the global latency budget across [region_footprint]; and failure modes including gateway loss, store partition, and backpressure handling.

Pro tips

  • Pin down the deliverable-3 crossover threshold early — the group size at which you switch from write fan-out to read fan-out decides your whole storage shape.
  • Set [concurrent_users] to a realistic peak, since it drives the gateway sharding and connection-to-node mapping math; under-sizing here cascades into every later number.
  • Match [connection_protocol] to your constraints; WebSocket is the default, but confirm the fallback path (SSE or long-poll) for clients that can't hold a socket.
  • Choose [message_store] for both write throughput and search needs, because chat history is write-heavy and read receipts add their own load.
  • Treat the [region_footprint] latency budget as real, and ask the AI to show where cross-region hops eat into it.
  • Don't skip backpressure handling; a design with no answer for a slow consumer or a store partition looks complete but fails under load.

Frequently Asked Questions

How does it handle 10M concurrent connections?
Through a connection layer with gateway sharding, sticky routing, and an explicit connection-to-node mapping sized for `[concurrent_users]`. No single node holds all sockets, so the design distributes live connections across many gateways with back-of-envelope math to justify the count.
Does it choose between at-least-once and exactly-once delivery?
Yes. Deliverable 2 covers delivery guarantees, ordering, and dedup keys. Exactly-once is expensive and often impractical, so the common approach is at-least-once delivery plus dedup keys on the client, which the design weighs explicitly rather than assuming.
What is the group fan-out crossover threshold?
It is the group size at which the design switches from write fan-out (pushing each message to every member's inbox) to read fan-out (members pull from a shared store). Pin this down early, because naive write fan-out to large groups melts the queue and it dictates your storage shape.
Does it address offline users and read receipts?
Yes, deliverable 4 includes offline message queueing so users receive messages sent while disconnected, plus read receipts. Both add load to the `[message_store]`, which is why the store choice should account for write throughput and search, not just capacity.
Is it useful for a system design interview?
Yes. Real-time chat at scale is a canonical interview topic, and the deliverables map to what reviewers probe: connection sharding, delivery guarantees, fan-out, and failure modes. Ask the AI to justify the capacity math so you can defend each number under follow-up questions.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in System Design Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support