Skip to main content

Claude/ChatGPT Prompt to Build an LLM Chatbot Eval Rubric and Dataset

Design an evaluation harness with a scored rubric and 50 diverse test cases for an LLM chatbot, including adversarial and out-of-scope prompts.

Vul de plaatshouders in

Edit the values, then copy your finished prompt.

Jouw Prompt
prompt.txt

                                

What this prompt does

This prompt builds an LLM chatbot evaluation harness concrete enough to run today — a scored rubric plus 50 diverse test cases, not abstract quality talk. It casts the AI as an ML engineer building evaluation. You provide the [use_case], the [forbidden] must-never behaviors, the required [format], and the [tone] target. It returns a five-dimension rubric (correctness, helpfulness, safety, tone, format adherence) with a 1–5 scale and per-point descriptors, 50 test prompts across happy path, edge cases, ambiguous intent, adversarial, and out-of-scope, expected behavior per prompt including correct refusals, a scoring guide for marginal answers, and a pass threshold.

The value lives in the adversarial and out-of-scope cases, which is where bots leak data or overreach. By deriving refusals directly from your [forbidden] list, the harness tests the exact failures that cause incidents. The per-point rubric descriptors and the marginal-answer scoring guide are what let two reviewers agree, so scores stay consistent across people and across prompt revisions.

When to use it

  • Standing up evaluation for a chatbot before shipping or after a prompt change
  • Writing refusal tests from a concrete [forbidden] behavior list
  • Building a reproducible rubric so two reviewers score consistently
  • Catching regressions in safety when you tweak a system prompt
  • Covering adversarial and out-of-scope inputs, not just the happy path
  • Setting a clear pass threshold that decides what blocks release

Example output

You get the rubric table first — five dimensions, each with a 1–5 scale and one-line descriptors per point — followed by 50 cases listed as (prompt, category, expected behavior). The cases span happy path, edge, ambiguous, adversarial, and out-of-scope, with correct refusals spelled out for the forbidden behaviors. A scoring guide for marginal answers and a pass threshold tell you which failures block release.

Pro tips

  • Make [forbidden] exhaustive and specific — "give refunds outside policy, share another customer's order, invent shipping dates" generates sharper refusal tests than a vague "be safe"
  • Pin the [format] exactly (e.g. short answer plus a single next-step link) so the format-adherence dimension has a concrete target
  • Write the refusals before the happy path — the adversarial and out-of-scope cases are where incidents start
  • Use the marginal-answer scoring guide when two reviewers disagree; it is built to resolve exactly those ties
  • Set the pass threshold strictly for safety failures, since a single leaked-order case should block release regardless of overall score
  • Re-run the 50 cases after every prompt change to catch silent safety regressions before users do
  • Grow the adversarial and out-of-scope buckets over time as you discover new ways users try to push the bot off its [use_case]

Frequently Asked Questions

Why does it focus on adversarial and out-of-scope cases?
Those are where bots leak data or overreach, so they catch the failures that actually cause incidents. The happy path rarely surprises you; the adversarial and out-of-scope prompts test whether the bot refuses correctly under pressure, which is where real risk concentrates.
How does the rubric keep two reviewers consistent?
Each of the five dimensions has a 1–5 scale with one-line descriptors per point, plus a separate scoring guide for marginal answers. That structure is designed so two reviewers reading the same answer land on the same score instead of relying on personal judgment.
Where do the refusal tests come from?
They are derived from your `[forbidden]` list, so the expected behavior for those cases is a correct refusal. The more specific that list is, the more precise the refusal tests become, which is why naming exact prohibited actions matters more than a generic safety goal.
Are 50 cases enough to trust a release?
Fifty diverse cases across five categories is a solid starting harness, not an exhaustive guarantee. It catches obvious regressions and the highest-risk failures, but for a high-stakes bot you will want to grow the set over time, especially the adversarial and out-of-scope buckets.
Engr Mejba Ahmed

Need this built for real?

Engr Mejba Ahmed

AI Developer · Software Engineer

I'm Mejba — I design and ship production AI systems, automations, and full-stack apps. If you want this turned into a working solution for your team, let's talk.

Meer in AI & Machine Learning Prompts

Engr Mejba Ahmed

Engr Mejba Ahmed

Claude Code Expert · Online

👋

Hey there!

Quick Actions

WhatsApp Instant reply

Chat on WhatsApp

+880 1723 741224 · Instant reply

Popular Questions

Engr Mejba Ahmed is connected
Engr Mejba Ahmed is typing...
Engr Mejba Ahmed avatar

✉ Want me to follow up? Drop your email

Engr Mejba Ahmed avatar

📞 Connect Directly

Choose how you'd like to reach me

WhatsApp

+880 1723 741224

Email

[email protected]

✓ Details sent! I'll get back to you shortly.

Powered by OpenAI

335+

Blog Posts

25

AI Courses

63

Projects

Services & Expertise

Pricing & Process

Learning & Resources

Connect & Support