What this prompt does
This prompt builds an LLM evaluation framework for a stated [evaluation_purpose]. It creates an evaluation dataset of [dataset_size] examples across [categories], automated metrics ([auto_metrics]), LLM-as-judge scoring on [judge_criteria], a human-evaluation rubric, A/B comparison between [models_to_compare], per-model cost and latency tracking, a regression suite for prompt changes, safety and toxicity checks, hallucination detection for [factual_domain], and a results dashboard, implemented in [language] with [eval_framework], including significance testing and alerts below [quality_threshold].
The structure works because it makes prompt changes measurable. Without a golden dataset and regression tests, you cannot tell whether a tweak helped or quietly broke something. LLM-as-judge scales evaluation beyond what humans can hand-grade, while statistical significance testing and confidence intervals stop you from over-reading noise as improvement. The cost-and-latency tracking keeps quality honest about its price, since the best-scoring model is not always the one worth shipping.
When to use it
- Before shipping any LLM feature where a prompt regression would hurt users.
- Comparing
[models_to_compare]on your own task rather than public benchmarks. - Catching silent quality drops when you edit a prompt or swap a model.
- Measuring whether a change actually improved output or just changed it.
- Tracking cost and latency alongside quality so you ship a viable, not just accurate, model.
- Detecting hallucinations in a
[factual_domain]where wrong answers are costly.
Example output
You get an evaluation harness: a labeled dataset of [dataset_size] examples spanning [categories], scripts computing [auto_metrics], an LLM-as-judge module scoring [judge_criteria], a human-rubric interface, an A/B comparison report across [models_to_compare], cost/latency logging, a regression test suite, safety and hallucination checks for [factual_domain], and a dashboard tracking scores over time with significance tests and threshold alerts at [quality_threshold], built in [language] with [eval_framework].
Pro tips
- Invest in the golden dataset;
[dataset_size]examples that truly span[categories]matter more than any clever metric. - Treat
[auto_metrics]like BLEU and ROUGE as weak signals; LLM-as-judge on[judge_criteria]usually correlates better with quality. - Always report confidence intervals; a 0.02 score bump inside the noise band is not a real improvement.
- Calibrate the LLM judge against a sample of human labels so you trust its scores before scaling.
- Keep the regression suite in CI so a prompt edit that drops below
[quality_threshold]fails loudly. - Log cost and latency next to quality; a marginally better model that is twice the price may not be worth shipping.