What this prompt does
This prompt designs a full evaluation framework for testing and comparing prompts — the harness that stops a prompt tweak from silently regressing in production. It produces [test_case_count] test cases with inputs, expected outputs, and scoring rubrics, then layers on weighted evaluation criteria, automated scoring functions (exact match, semantic similarity, regex, and LLM-as-judge), an A/B design across [variants], a regression suite, and an edge-case suite covering empty, oversized, adversarial, and multilingual inputs. It also tracks cost and latency, computes statistical significance, and outlines version control and a results dashboard.
The variables shape what gets tested and how. [prompt_purpose] defines the kind of prompt under evaluation, and [criteria] sets the weighted dimensions — correctness, completeness, code quality, and clarity, for example — that drive scoring. [variants] names the prompt versions to compare in the A/B test, while [language] and [eval_framework] (such as promptfoo) decide the implementation. Together they turn a subjective "this looks better" into a measurable, repeatable evaluation.
When to use it
- You are about to ship an LLM feature and want a guardrail against silent prompt regressions.
- A prompt change looked fine on one example and you need confidence it holds across ten or twenty.
- You are comparing prompt variants (
[variants]) and want statistically grounded A/B results, not gut feel. - Edge cases — empty, very long, adversarial, or multilingual input — keep slipping through and need their own suite.
- You need cost and latency tracking per variant to balance quality against budget.
- Establishing prompt version control with a changelog so changes are auditable.
Example output
Expect a framework spec and scaffolding: a table of [test_case_count] test cases with inputs, expected outputs, and rubrics; scoring functions mapped to each criterion in [criteria]; an A/B comparison design across [variants] with a significance calculation; separate regression and edge-case suites; and a report template showing pass rates and trends. If you supply [eval_framework], it is implemented in [language] against that tool rather than from scratch.
Pro tips
- Weight
[criteria]to reflect what actually matters; if correctness is 40 percent, the framework will rank a stylish-but-wrong answer below a plain correct one, which is usually what you want. - Build the regression suite from prompts that previously worked, since catching a change that breaks known-good behavior is the whole point.
- Use LLM-as-judge for fuzzy criteria like clarity, but pin exact match or regex on anything that must be precise, so scoring is not all subjective.
- Make the edge-case suite genuinely adversarial — empty input, oversized input, and injection-style input are where prompts quietly fail.
- Track cost and latency per variant alongside quality; the best-scoring prompt is not worth shipping if it is too slow or expensive at your
[variants]volume. - Keep prompts under version control with a changelog so you can trace which edit moved which metric.