What this prompt does
This prompt takes the seven concrete variables that actually determine model fit — task type, dataset size, feature composition, label distribution, latency budget, infrastructure, and team experience — and uses them to drive a structured recommendation rather than a generic "try XGBoost" answer. The [label_distribution] and [data_issues] fields are what separate useful output from textbook advice: they force the model to reason about class imbalance, missing-value strategies, and data leakage before it ever suggests an architecture.
The template produces a ranked shortlist of three models, a deliberate baseline, feature engineering direction, a training and tuning strategy, evaluation thresholds, known pitfalls, and a code scaffold — all in one pass. That last point matters: you do not need to run eight separate conversations to get from problem statement to working code skeleton.
When to use it
- You are starting a new supervised learning project and need to justify your model choice to stakeholders before spending sprint time on experiments.
- Your dataset is in an unusual shape — very wide (1000+ features, few rows), heavily imbalanced, or has categorical columns with high cardinality — and you want guidance specific to that shape.
- A previous model is underperforming in production and you want a structured audit of whether the architecture choice was the problem.
- You are an ML engineer onboarding a new team member and want a concrete starting point with documented trade-offs.
- You need to match a latency constraint (sub-50ms inference on CPU) and want model options ranked by that constraint alongside accuracy.
- You are deciding between training from scratch versus fine-tuning a pretrained model and want the trade-off written out explicitly.
Example output
For [ml_task_type]: binary classification, [business_problem]: churn prediction for a SaaS product, [dataset_size]: 80,000 rows, [label_distribution]: 8% positive (churned):
Top 3 recommendations:
1. LightGBM (primary) — handles imbalance via scale_pos_weight, fast iteration,
interpretable via SHAP. Con: needs careful max_depth tuning to avoid overfit
on low-signal behavioral features.
2. Logistic Regression (with L2) — useful calibration baseline; probabilities
are reliable out-of-the-box for threshold tuning. Con: misses feature interactions.
3. CatBoost — strong on mixed categorical/numeric without extensive preprocessing.
Con: slower training, higher memory footprint than LGBM.
Baseline: Logistic Regression on top-10 features by mutual information.
Expected AUC: 0.72-0.76. Use this to benchmark everything else.
Primary metric: ROC-AUC. Secondary: Precision-Recall AUC (more informative
under 8% positive rate). Production threshold: PR-AUC > 0.42 before ship.
Common pitfall: leaking post-churn activity features (e.g. last_login computed
after label window closes). Audit feature timestamps against label cutoff.
Pro tips
- Fill
[data_issues]honestly. If you write "none," you will get a generic answer. Writing "15% nulls in usage columns, two features with train/test distribution shift" unlocks specific imputation and drift-detection advice that changes the training strategy. - Set
[latency_requirement]in milliseconds, not words. "Fast" means nothing. "Less than 30ms p99 on a single CPU core" eliminates ensemble methods and tree-boosting and steers the output toward linear models or quantized neural nets. - Use
[team_experience]to control complexity. "Junior team, no MLOps infra" should produce a simpler pipeline recommendation than "senior team, Kubeflow available." The prompt respects this — do not downplay it to seem more capable. - Run the baseline code first, always. The prompt outputs a code starter for the baseline model. Run that before touching the top recommendation. If the baseline hits your metric threshold, you are done — ship it.
- Pair this with a data profiling step. Before filling
[label_distribution]and[data_issues], rundf.describe(),df.isnull().sum(), and a quick value-count on your target column. Five minutes of profiling makes every field more accurate and the output significantly more actionable.