What this prompt does
This prompt turns a vague "the prompt seems fine" feeling into a repeatable evaluation harness. It instructs the model to act as a senior prompt engineer and produce six concrete artifacts: a curated golden dataset, a scoring rubric, an LLM-as-judge setup, a versioned results table, a regression gate, and runnable code that ties it all together. The framing — "a harness a team can run on every prompt change, not a one-off spot check" — pushes the output toward something durable rather than a throwaway test.
The four context variables shape every artifact. [task] defines what the prompt under test actually does, so the golden inputs are realistic. [dimensions] becomes the rubric axes the judge scores against, so you grade what matters instead of generic "quality." [acceptance_rule] is the objective bar each golden output must clear, which keeps the expected outputs honest. [regression_gate] is the pass/fail threshold that decides whether a prompt change ships — wiring it in early means the harness fails loudly instead of drifting silently.
When to use it
- You maintain an LLM step inside a product and want to stop guessing whether edits broke something
- You're about to refactor a long prompt and need a safety net before you touch it
- You want to compare two prompt versions on the same inputs with a numeric verdict
- You're setting up CI for prompts and need a gate that blocks regressions automatically
- A teammate keeps "improving" prompts by feel and you want objective scoring
- You're onboarding to an existing prompt and want a baseline score to measure against
Example output
You get runnable harness code first — a script that loads the golden inputs, calls the prompt, runs each output through the judge, and prints a table. After that comes a sample rubric (each dimension scored 1-5 with written descriptors) and a sample results table showing scores per prompt version, so regressions are visible at a glance. The regression gate is stated as a concrete threshold, not a vibe.
Pro tips
- Keep the golden set small but adversarial — set
[task]precisely and choose 30-50 inputs where a few hard edge cases catch most regressions - Make
[dimensions]specific and measurable; "faithfulness, coverage, brevity" judges far better than a single "quality" axis - Write
[acceptance_rule]as something a judge can check mechanically (e.g. word count, required elements) so scoring is consistent - Set
[regression_gate]to fail on both a mean-score drop and any single catastrophic case scoring 1, so one bad output can't hide behind a good average - Pin the judge model and temperature; a drifting judge makes the whole table untrustworthy
- Re-run the harness on every prompt edit, not just before release, so you catch regressions while the change is fresh in your head