What this prompt does
This prompt makes the AI a machine learning systems architect designing a recommendation engine for a [platform_type] platform with [user_count] users and [item_count] items. It covers requirements, a data-collection pipeline, candidate generation, ranking, the serving architecture, and an experimentation framework. The core is a two-stage design: cheap candidate generation via collaborative filtering and content-based signals, then expensive ranking with [ranking_model] over a merged pool of up to [candidate_pool] items.
The structure works because separating recall from scoring keeps latency sane. Candidate generation uses matrix factorization to produce [embedding_dim]-dimensional user and item embeddings and approximate nearest-neighbor search via [ann_index] to pull similar items fast, while ranking applies a heavier model only to that shortlist to meet [serving_latency]. Running collaborative filtering and content-based matching in parallel gives the system recall even for new users or items where one signal alone is weak. The serving layer pre-computes batch recommendations in [serving_store], refreshed every [batch_interval], and merges them with live results, while a re-ranking layer applies business rules and the A/B layer hashes users into [experiment_buckets] buckets to measure what actually moves the metric.
When to use it
- You're adding personalization and want the two-stage recall-then-rank pattern designed properly
- You need a data pipeline capturing both implicit and explicit signals
- You're choosing between collaborative filtering and content-based candidate generation
- You want a ranking model trained against a real
[optimization_metric]rather than raw clicks - You need a serving design that hits a strict
[serving_latency]budget - You want an A/B framework with significance testing and a multi-armed bandit built in
- You're serving a large
[item_count]catalog where exact nearest-neighbor search won't scale
Example output
Expect a layered ML design: a data-collection pipeline ingesting implicit and explicit signals through [stream_processor] into [data_lake], a candidate-generation section with matrix-factorization embeddings and [ann_index] retrieval plus content-based feature matching, a ranking stage listing the scoring features and the [optimization_metric] target, a serving architecture combining pre-computed lists in [serving_store] with live candidates and a re-ranking layer for business rules like boosting promoted items, and an experimentation framework with bucketing, per-bucket metrics, statistical significance, and a multi-armed bandit. It's a reasoned architecture, not code.
Pro tips
- Keep
[candidate_pool]bounded; the whole point of two stages is to rank a shortlist, not the full[item_count]catalog - Match
[embedding_dim]to your data volume — larger embeddings capture more but cost memory and ANN-search time - Choose
[optimization_metric]to reflect business value, not just clicks; optimizing the wrong target quietly misleads the model - Use
[ann_index]for fast approximate recall; exact nearest-neighbor search won't scale to[item_count]items - Pre-compute batch recs in
[serving_store]and merge live results, rather than computing everything per request - Treat the A/B layer as first-class; without
[experiment_buckets]and significance testing you can't tell if changes actually help