Skip to main content

Claude/ChatGPT Prompt to Brainstorm Features for Tabular ML Data

Brainstorm features for tabular ML data: generate 25 candidate features, each with a falsifiable hypothesis, leakage risk, and expected lift you can test.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt

                                

What this prompt does

This prompt does feature engineering for tabular data, generating 25 candidate features a team could compute and test — each with a falsifiable hypothesis, not generic ideas. It casts the AI as a senior ML engineer. You provide your [columns], the [target], the [time_horizon], and known [constraints]. It returns 25 features spread across temporal, aggregation, interaction, ratio, categorical-encoding, text, and external categories, each with a name, definition, and a hypothesis for why it predicts the target, plus a leakage-risk flag, an expected-lift ranking, and the three to test first.

Tying every feature to a falsifiable hypothesis is what separates this from a brainstorm dump — each idea states why it should predict the [target] so you can actually test it. The leakage-risk flag respects your [time_horizon], catching the classic trap of using data computed after the prediction moment. Ranking by expected lift and naming three to test first keeps you from building all 25 when most lift comes from a few.

When to use it

  • Engineering features for a tabular model from a known set of [columns]
  • Generating hypotheses you can falsify rather than a vague feature wishlist
  • Guarding against leakage by respecting a strict [time_horizon]
  • Prioritizing which features to build first by expected lift
  • Honoring data [constraints] like no-PII or nullable columns
  • Covering feature categories you might not brainstorm alone (text, external, interactions)

Example output

You get a table with columns for feature, definition, hypothesis, leakage risk, and expected lift, spanning all seven feature categories. Below it, the top three features to test first are called out with reasoning. Each row is specific enough to compute, and the leakage flags are tied to your stated time horizon so risky features are obvious at a glance.

Pro tips

  • Spell out the [time_horizon] precisely — "features must use data available before the 30-day window" is what makes the leakage flags accurate
  • List [columns] with their meanings, not just names, so the hypotheses are grounded in what the data actually represents
  • State [constraints] like no PII features or country may be null for 8% of rows so flagged features respect them
  • Build the top three first and measure lift before constructing the full 25; most of the signal usually lives in a few features
  • Scrutinize every feature flagged for leakage — anything computed after the prediction moment will inflate offline scores and fail in production
  • Re-run with a different [target] to reuse the same columns for a second model without starting from scratch
  • Use the seven feature categories as a checklist; the text and external buckets are the ones engineers most often forget to brainstorm

Frequently Asked Questions

Why does each feature need a hypothesis?
A falsifiable hypothesis states why the feature should predict the `[target]`, which makes it testable rather than a guess. That discipline separates the output from a generic brainstorm and lets you discard features whose stated reasoning does not hold up against the data.
How does it handle data leakage?
Every feature carries a leakage-risk flag tied to your stated `[time_horizon]`. The classic trap is computing a feature from data that only exists after the prediction moment, so a precise time boundary is what lets the prompt flag those risky features reliably.
Do I have to build all 25 features?
No — the output ranks features by expected lift and names three to test first. Most lift usually comes from a few features, so building and measuring the top three before constructing the rest saves effort and avoids over-engineering the feature set.
Can it respect constraints like no PII?
Yes — state them in `[constraints]` and it notes where a feature would violate one or needs the listed external data. For example, with `no PII features` it avoids identity-based features, and with nullable columns it flags features that depend on the missing values.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in AI & Machine Learning Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

Claude Code Expert · Online

👋

Hey there!

Quick Actions

WhatsApp Instant reply

Chat on WhatsApp

+880 1723 741224 · Instant reply

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

[email protected]

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support