Skip to main content

Claude/ChatGPT Prompt to Build a Great Expectations Data Quality Suite

Generate a production Great Expectations test suite for a critical table: schema, nulls, ranges, FK checks, and drift detection.

Fill in the placeholders

Edit the values, then copy your finished prompt.

Your Prompt
prompt.txt
You are a senior data engineer. Specify a Great Expectations test suite tightly enough to drop into CI, not a vague checklist.

Context:
- Table: fct_orders
- Warehouse: Snowflake
- Primary key: order_id
- Freshness window: last 24 hours

Deliverables:
1. Schema checks: expected columns, types, and ordering.
2. Non-null expectations on every business-critical column.
3. Uniqueness on the primary key and any natural keys.
4. Value-range and allowed-set expectations for numeric and enum columns.
5. Foreign-key existence checks against parent tables plus a row-count-vs-prior-run guard.
6. String regex patterns and a distribution-drift expectation tied to the freshness window.

Output: the JSON expectation suite, plus a one-line doc note explaining what each expectation protects against.

What this prompt does

This prompt asks the model to specify a Great Expectations test suite tight enough to drop straight into CI, not a vague checklist. You provide [table], [warehouse], [primary_key], and [freshness], and it returns a JSON expectation suite covering schema, nulls, uniqueness, value ranges, foreign keys, regex patterns, and drift -- each expectation annotated with what it protects against.

The structure works because it maps each common data-quality failure to a specific expectation. [primary_key] drives the uniqueness checks and the foreign-key existence tests against parent tables. [freshness] anchors the row-count-vs-prior-run guard and the distribution-drift expectation, so the suite knows what "recent" means. [table] and [warehouse] ground the schema and type checks in the real columns. By demanding a one-line note per expectation, the prompt produces a suite a reviewer can reason about instead of a wall of opaque rules, which is what keeps the checks maintained over time rather than ignored.

When to use it

  • A table feeds something a customer sees and you want quality gates before it ships.
  • You need schema, null, uniqueness, and range checks generated as a real GE suite.
  • You want foreign-key existence checks against parent tables written for you.
  • You need a row-count-vs-prior-run guard to catch partial or doubled loads.
  • You want distribution-drift detection tied to a freshness window.
  • You're wiring data quality into CI and need JSON you can commit, not a checklist.
  • You want regex pattern checks on formatted string columns to catch malformed values.

Example output

You get the JSON expectation suite -- schema and type checks, non-null expectations on business-critical columns, uniqueness on the primary key and natural keys, value-range and allowed-set rules, foreign-key checks, regex patterns, and a drift expectation -- with a one-line doc note beside each explaining what it guards against. It's structured to drop into a Great Expectations checkpoint and run in CI without you having to hand-author each rule from scratch.

Pro tips

  • Name the real [primary_key] (order_id) so the uniqueness and FK checks key on the right column.
  • Set [freshness] to your actual load cadence (last 24 hours) so the row-count guard and drift check use a sensible window.
  • Make [table] and [warehouse] specific so the schema and type expectations match what's really deployed.
  • Wire the suite into CI so a bad load fails loud instead of leaking downstream into a dashboard nobody re-checks.
  • Ask the model to flag which columns are business-critical so non-null checks land where they matter, not on every field.
  • Tune allowed-set and range expectations to your data; over-tight ranges cause false alarms that erode trust in the suite.
  • Add regex pattern checks on string columns that follow a known format so malformed values fail the suite early.

Frequently Asked Questions

Is the output ready to drop into CI?
Yes. The prompt explicitly asks for a suite tight enough to drop into CI, not a vague checklist, returned as a JSON expectation suite. Wiring it into CI means a bad load fails loud rather than leaking into a dashboard nobody double-checks.
Does it generate foreign-key checks?
It includes foreign-key existence checks against parent tables, keyed off the `[primary_key]` and natural keys you specify. It also adds a row-count-vs-prior-run guard, which together catch broken references and partial or doubled loads.
How does the drift detection work?
The suite includes a distribution-drift expectation tied to your `[freshness]` window, so it compares recent data against expected distribution over a window you define. Tune the ranges to your real data, since over-tight bounds produce false alarms that erode trust.
What do the one-line notes on each expectation do?
Each expectation comes with a one-line doc note explaining what it protects against. That makes the suite reviewable, so a teammate can reason about why a check exists rather than facing a wall of opaque rules with no rationale.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

More in Data Engineering & ETL Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

AI assistant · trained on my work

👋

Hey there!

Quick Actions

WhatsApp Direct line to me

Chat on WhatsApp

+880 1723 741224 · Replies within the hour on working days

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

mejba.13@gmail.com

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support