Skip to main content

System Design Prompt: Build a Video Streaming Platform

Design a video streaming platform: adaptive bitrate, transcoding pipeline, multi-tier CDN, live streaming, and a recommendation engine.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a senior distributed systems engineer. Design a video streaming platform similar to YouTube that serves 10 million concurrent viewers.

Step 1: Define the core user flows: video upload and processing, video discovery and search, adaptive bitrate playback, live streaming, and user engagement (likes, comments, watch history). Specify non-functional requirements: startup latency under 2 seconds, rebuffer ratio below 1%, support for 4K maximum resolution, and global availability across 6 regions.

Step 2: Design the video ingestion and transcoding pipeline. When a creator uploads a video, store the raw file in Amazon S3. Trigger a transcoding job queue that produces 6 bitrate variants (from 360p to 4K) in HLS and DASH formats. Generate thumbnails at key frames and extract metadata (duration, codec, aspect ratio). Design for processing 500,000 uploads per day.

Step 3: Architect the content delivery layer using a multi-tier CDN strategy. Place origin servers in 6 regions, edge caches at ISP peering points, and a CloudFront CDN for global distribution. Implement cache warming for trending content and predictive pre-fetching based on user watch patterns. Design the cache eviction policy to maximize the hit ratio for the content catalog.

Step 4: Design the adaptive bitrate streaming logic. The player client should measure available bandwidth every 2 seconds and select the optimal quality level. Implement buffer management with a target buffer length that balances startup time against rebuffering risk. Handle network transitions (WiFi to cellular) gracefully.

Step 5: Build the recommendation engine architecture using collaborative filtering and content-based signals. Ingest watch history, completion rates, explicit ratings, and search queries. Design a two-stage pipeline: candidate generation (retrieves top 500 from Redis + Elasticsearch) followed by ranking (scores candidates using a trained model). Serve recommendations with under 100ms latency.

Step 6: Design the live streaming architecture that supports 10,000 simultaneous broadcasts. Implement the ingest protocol (RTMP/SRT), real-time transcoding to adaptive bitrate, edge distribution with under 5 seconds glass-to-glass latency, and a chat system that scales with viewer count.

What this prompt does

This prompt makes the AI design a video streaming platform similar to [reference_platform] that serves [concurrent_viewers] concurrent viewers. It works through six layers: core user flows and non-functional targets (startup under [startup_latency], up to [resolution], across [region_count] regions), a transcoding pipeline producing [bitrate_count] variants from raw files in [object_storage], a multi-tier CDN on [cdn_provider], adaptive bitrate playback logic, a two-stage recommendation engine, and a live-streaming architecture for [concurrent_live_streams] broadcasts.

The structure works because streaming is fundamentally a storage-and-bandwidth problem solved in stages. Splitting transcoding from delivery from playback lets each layer scale independently: the pipeline absorbs [daily_uploads] uploads a day from raw files in [object_storage], the CDN warms trending content at the edge and pre-fetches based on watch patterns, and the player measures bandwidth every [measurement_interval] to pick a quality level without rebuffering. The recommendation step deliberately separates cheap candidate generation from a ranking model, which keeps serving latency within [recommendation_latency] even against a large catalog. Designing across [region_count] regions up front also means the delivery and origin layout reflects where your viewers actually are.

When to use it

  • You're structuring an ingestion, delivery, and recommendation stack before sizing real infrastructure
  • You need to reason about transcoding fan-out into multiple bitrates and formats
  • You're planning a multi-tier CDN strategy with cache warming and predictive pre-fetching
  • You want adaptive bitrate logic that handles WiFi-to-cellular transitions gracefully
  • You're adding live streaming with a glass-to-glass [live_latency] target
  • You need a recommendation architecture that won't blow past your serving budget
  • You want startup latency, rebuffer ratio, and max [resolution] pinned as explicit targets

Example output

Expect a layered design document: the upload-to-playback flow, a transcoding pipeline producing [bitrate_count] HLS and DASH renditions with key-frame thumbnails and extracted metadata, a CDN topology with origin regions and edge caches plus a cache-eviction policy, the client-side ABR algorithm and buffer-management approach that balances startup time against rebuffering, a two-stage candidate-generation-then-ranking recommendation pipeline, and a live-streaming path covering RTMP/SRT ingest, real-time transcoding, edge distribution, and chat that scales with viewer count. It's an architecture walkthrough with reasoning at each layer, not code.

Pro tips

  • Set [concurrent_viewers] and [concurrent_live_streams] to realistic peaks; they drive CDN and ingest sizing more than anything else
  • Match [bitrate_count] to your audience — more renditions smooth playback but multiply transcoding cost and storage
  • Keep [startup_latency] and the live [live_latency] target separate; pre-recorded and live have very different buffering tradeoffs
  • Tune the ABR [measurement_interval] carefully: too frequent and you flap between qualities, too slow and you rebuffer on network drops
  • Use [object_storage] for raw uploads and let the CDN serve renditions; serving directly from object storage gets expensive fast
  • The recommendation two-stage split matters most when the catalog is large — don't rank [item_count]-scale candidates directly

Frequently Asked Questions

Why produce multiple bitrate variants instead of one high-quality stream?
Multiple renditions enable adaptive bitrate streaming, where the player picks the best quality the current connection can sustain. Producing `[bitrate_count]` variants from 360p up to `[resolution]` lets viewers on weak networks keep watching at lower quality instead of rebuffering, which keeps the rebuffer ratio low.
What's the point of separating candidate generation from ranking in recommendations?
Candidate generation cheaply narrows millions of items to a few hundred, and ranking then scores only that small set with a heavier model. This two-stage split keeps recommendation serving within `[recommendation_latency]`, which you couldn't hit by running an expensive model over the whole catalog.
Does this cover live streaming or only on-demand video?
It covers both. Step six designs the live path specifically, including RTMP/SRT ingest, real-time transcoding to adaptive bitrate, edge distribution targeting `[live_latency]` glass-to-glass, and a chat system that scales with viewer count, alongside the on-demand pipeline in the earlier steps.
Is the design locked to a particular CDN or cloud provider?
No. It defaults to `[cdn_provider]` and `[object_storage]` values but the multi-tier strategy is provider-neutral. You set those variables to your stack, and the origin, edge-cache, and cache-warming concepts carry over regardless of which CDN or object store you choose.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in System Design Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support