Skip to main content

Design a Recommendation Engine

Design a scalable recommendation engine with collaborative filtering, content-based signals, real-time personalization, and A/B testing.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a machine learning systems architect. Design a recommendation engine for a e-commerce marketplace platform serving 50 million users with a catalog of 10 million items.

Step 1: Define the recommendation requirements. The system must generate personalized recommendations for the home feed (home feed, search results, item detail pages, and email digests). Target a click-through rate improvement of 30% over the non-personalized baseline. Support both real-time recommendations (respond within 50ms) and batch-generated recommendations refreshed every 6 hours.

Step 2: Design the data collection pipeline. Ingest implicit signals (views, clicks, time spent, scroll depth, add-to-cart, purchases) and explicit signals (ratings, likes, saves, reviews) from Kafka topic. Process events through Apache Flink for real-time feature updates. Store the raw event log in S3 + Delta Lake for batch model training. Define the user-item interaction matrix schema and estimate its size.

Step 3: Implement the candidate generation stage using two parallel strategies. First, collaborative filtering: train a matrix factorization model (ALS) on the user-item interaction matrix to generate 128-dimensional user and item embeddings. Use approximate nearest neighbor search via FAISS to find the top 200 similar items. Second, content-based filtering: extract item features (category, attributes, text embeddings from descriptions) and match against the user's preference profile.

Step 4: Design the ranking stage that takes the merged candidate set (up to 500 items) and scores each using a gradient-boosted decision tree (XGBoost). Features include: candidate generation score, user-item affinity, item popularity, item freshness, user context (time of day, device, location), and diversity penalty. Train the ranking model on historical engagement data with the optimization target being expected revenue per impression.

Step 5: Build the serving architecture. Pre-compute batch recommendations for all users and store in Redis Cluster. For real-time requests, merge the pre-computed list with live candidate generation results. Implement a re-ranking layer that applies business rules (boost promoted items, filter blocked items, enforce diversity constraints). Serve the final list within 50ms.

Step 6: Design the experimentation framework for A/B testing recommendation algorithms. Hash users into 100 experiment buckets. Track key metrics (CTR, conversion rate, engagement time, revenue per user) per bucket. Implement a multi-armed bandit for automatic traffic allocation to winning variants. Build a dashboard that shows statistical significance and confidence intervals.

What this prompt does

This prompt makes the AI a machine learning systems architect designing a recommendation engine for a [platform_type] platform with [user_count] users and [item_count] items. It covers requirements, a data-collection pipeline, candidate generation, ranking, the serving architecture, and an experimentation framework. The core is a two-stage design: cheap candidate generation via collaborative filtering and content-based signals, then expensive ranking with [ranking_model] over a merged pool of up to [candidate_pool] items.

The structure works because separating recall from scoring keeps latency sane. Candidate generation uses matrix factorization to produce [embedding_dim]-dimensional user and item embeddings and approximate nearest-neighbor search via [ann_index] to pull similar items fast, while ranking applies a heavier model only to that shortlist to meet [serving_latency]. Running collaborative filtering and content-based matching in parallel gives the system recall even for new users or items where one signal alone is weak. The serving layer pre-computes batch recommendations in [serving_store], refreshed every [batch_interval], and merges them with live results, while a re-ranking layer applies business rules and the A/B layer hashes users into [experiment_buckets] buckets to measure what actually moves the metric.

When to use it

  • You're adding personalization and want the two-stage recall-then-rank pattern designed properly
  • You need a data pipeline capturing both implicit and explicit signals
  • You're choosing between collaborative filtering and content-based candidate generation
  • You want a ranking model trained against a real [optimization_metric] rather than raw clicks
  • You need a serving design that hits a strict [serving_latency] budget
  • You want an A/B framework with significance testing and a multi-armed bandit built in
  • You're serving a large [item_count] catalog where exact nearest-neighbor search won't scale

Example output

Expect a layered ML design: a data-collection pipeline ingesting implicit and explicit signals through [stream_processor] into [data_lake], a candidate-generation section with matrix-factorization embeddings and [ann_index] retrieval plus content-based feature matching, a ranking stage listing the scoring features and the [optimization_metric] target, a serving architecture combining pre-computed lists in [serving_store] with live candidates and a re-ranking layer for business rules like boosting promoted items, and an experimentation framework with bucketing, per-bucket metrics, statistical significance, and a multi-armed bandit. It's a reasoned architecture, not code.

Pro tips

  • Keep [candidate_pool] bounded; the whole point of two stages is to rank a shortlist, not the full [item_count] catalog
  • Match [embedding_dim] to your data volume — larger embeddings capture more but cost memory and ANN-search time
  • Choose [optimization_metric] to reflect business value, not just clicks; optimizing the wrong target quietly misleads the model
  • Use [ann_index] for fast approximate recall; exact nearest-neighbor search won't scale to [item_count] items
  • Pre-compute batch recs in [serving_store] and merge live results, rather than computing everything per request
  • Treat the A/B layer as first-class; without [experiment_buckets] and significance testing you can't tell if changes actually help

Frequently Asked Questions

Why split recommendations into candidate generation and ranking?
Candidate generation cheaply narrows `[item_count]` items down to a shortlist using embeddings and approximate nearest-neighbor search, and ranking then scores only that pool with a heavier model. This two-stage split is what keeps serving within `[serving_latency]`, since ranking the full catalog per request would be far too slow.
What's the difference between collaborative and content-based filtering here?
Collaborative filtering learns from user-item interaction patterns via matrix factorization, recommending items that similar users engaged with. Content-based filtering matches item features like category and text embeddings against a user's preference profile. The prompt runs both in parallel during candidate generation so you get recall even for new users or items.
How does the A/B testing framework decide which algorithm wins?
Users are hashed into `[experiment_buckets]` buckets, and metrics like click-through, conversion, and revenue per user are tracked per bucket with statistical significance and confidence intervals. A multi-armed bandit can then shift traffic toward winning variants automatically rather than waiting for a manual call.
Can this serve both real-time and batch recommendations?
Yes. Batch recommendations are pre-computed for all users and stored in `[serving_store]`, refreshed every `[batch_interval]`. For real-time requests, the serving layer merges those pre-computed lists with live candidate-generation results and re-ranks, so you get freshness without recomputing everything on each call.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in System Design Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support