What this prompt does
This prompt makes the AI design a video streaming platform similar to [reference_platform] that serves [concurrent_viewers] concurrent viewers. It works through six layers: core user flows and non-functional targets (startup under [startup_latency], up to [resolution], across [region_count] regions), a transcoding pipeline producing [bitrate_count] variants from raw files in [object_storage], a multi-tier CDN on [cdn_provider], adaptive bitrate playback logic, a two-stage recommendation engine, and a live-streaming architecture for [concurrent_live_streams] broadcasts.
The structure works because streaming is fundamentally a storage-and-bandwidth problem solved in stages. Splitting transcoding from delivery from playback lets each layer scale independently: the pipeline absorbs [daily_uploads] uploads a day from raw files in [object_storage], the CDN warms trending content at the edge and pre-fetches based on watch patterns, and the player measures bandwidth every [measurement_interval] to pick a quality level without rebuffering. The recommendation step deliberately separates cheap candidate generation from a ranking model, which keeps serving latency within [recommendation_latency] even against a large catalog. Designing across [region_count] regions up front also means the delivery and origin layout reflects where your viewers actually are.
When to use it
- You're structuring an ingestion, delivery, and recommendation stack before sizing real infrastructure
- You need to reason about transcoding fan-out into multiple bitrates and formats
- You're planning a multi-tier CDN strategy with cache warming and predictive pre-fetching
- You want adaptive bitrate logic that handles WiFi-to-cellular transitions gracefully
- You're adding live streaming with a glass-to-glass
[live_latency]target - You need a recommendation architecture that won't blow past your serving budget
- You want startup latency, rebuffer ratio, and max
[resolution]pinned as explicit targets
Example output
Expect a layered design document: the upload-to-playback flow, a transcoding pipeline producing [bitrate_count] HLS and DASH renditions with key-frame thumbnails and extracted metadata, a CDN topology with origin regions and edge caches plus a cache-eviction policy, the client-side ABR algorithm and buffer-management approach that balances startup time against rebuffering, a two-stage candidate-generation-then-ranking recommendation pipeline, and a live-streaming path covering RTMP/SRT ingest, real-time transcoding, edge distribution, and chat that scales with viewer count. It's an architecture walkthrough with reasoning at each layer, not code.
Pro tips
- Set
[concurrent_viewers]and[concurrent_live_streams]to realistic peaks; they drive CDN and ingest sizing more than anything else - Match
[bitrate_count]to your audience — more renditions smooth playback but multiply transcoding cost and storage - Keep
[startup_latency]and the live[live_latency]target separate; pre-recorded and live have very different buffering tradeoffs - Tune the ABR
[measurement_interval]carefully: too frequent and you flap between qualities, too slow and you rebuffer on network drops - Use
[object_storage]for raw uploads and let the CDN serve renditions; serving directly from object storage gets expensive fast - The recommendation two-stage split matters most when the catalog is large — don't rank
[item_count]-scale candidates directly