What this prompt does
This prompt makes the AI design a real-time chat system end to end, specified tightly enough to build with concrete components, protocols, and numbers. You set the [concurrent_users], the [connection_protocol], the [message_store], and the [region_footprint]. It returns a connection layer with gateway sharding and sticky routing, message delivery guarantees with ordering and dedup keys, a group fan-out strategy with a crossover threshold, storage and search plus offline queueing and read receipts, presence and E2E encryption hooks with a latency budget, and failure modes.
The structure works because the part that bites in real-time systems is fan-out: naive write fan-out to large groups melts your queue. Forcing the connection sharding and store decisions before code pressure-tests the design. [concurrent_users] drives the gateway sharding and connection-to-node mapping math, [connection_protocol] sets the transport, [message_store] shapes storage and search, and [region_footprint] defines the global latency budget.
When to use it
- You're designing a chat or messaging feature and need the connection layer right.
- You're prepping a system design interview on real-time systems at scale.
- You need to reason about gateway sharding and sticky routing for many sockets.
- You want delivery guarantees, ordering, and dedup keys defined explicitly.
- You need the group fan-out crossover threshold (write versus read) pinned down.
- You want presence, offline queueing, read receipts, and named failure modes.
Example output
You get an architecture diagram description, a per-component table, and back-of-envelope capacity math: a connection layer with gateway sharding, sticky routing, and connection-to-node mapping for [concurrent_users] live sockets; message delivery guarantees (at-least-once versus exactly-once) with ordering and dedup keys; a group fan-out strategy (write versus read) with the crossover threshold; storage and search using [message_store] plus offline message queueing and read receipts; presence, E2E encryption hooks, and the global latency budget across [region_footprint]; and failure modes including gateway loss, store partition, and backpressure handling.
Pro tips
- Pin down the deliverable-3 crossover threshold early — the group size at which you switch from write fan-out to read fan-out decides your whole storage shape.
- Set
[concurrent_users]to a realistic peak, since it drives the gateway sharding and connection-to-node mapping math; under-sizing here cascades into every later number. - Match
[connection_protocol]to your constraints; WebSocket is the default, but confirm the fallback path (SSE or long-poll) for clients that can't hold a socket. - Choose
[message_store]for both write throughput and search needs, because chat history is write-heavy and read receipts add their own load. - Treat the
[region_footprint]latency budget as real, and ask the AI to show where cross-region hops eat into it. - Don't skip backpressure handling; a design with no answer for a slow consumer or a store partition looks complete but fails under load.