Skip to content

01 · Evaluating Prompts: Test Sets & Rubrics

An evaluation ("eval") is a repeatable way to measure how well a prompt does its job. It has three parts: a test set of inputs, a way of scoring each output, and a summary you can compare across prompt versions. Without evals, every prompt change is a guess; with them, you can make changes confidently and catch regressions.

Building a test set

A good test set is small enough to run often and varied enough to be meaningful.

Where cases come from:

  • Real inputs (anonymized): the best source, because they have real messiness.
  • Known failures: every bug report becomes a test case.
  • Deliberate edge cases: empty input, extremely long input, wrong language, input containing instructions, input where the correct answer is "I don't know".
  • Synthetic cases: model-generated variations can fill gaps, but review them — synthetic inputs tend to be cleaner and more uniform than real ones.

Structure it by slice. Tag each case (e.g. short, angry, non-English, missing-field), so you can see where a prompt fails, not only how often.

Size: start with 20–50 cases for a focused task. Grow it as you find failures. Fifty well-chosen cases beat five hundred near-duplicates.

Reference answers: for tasks with a right answer (classification, extraction), store the expected output. For open-ended tasks, store what a good answer must contain or must avoid.

Scoring: three kinds of check

Check Good for Example
Exact / programmatic Labels, JSON validity, fields, format, length, forbidden words output["category"] == expected
Reference-based Extraction, short answers Normalized string match; set overlap of extracted items
Rubric-based (human or model judge) Writing quality, helpfulness, faithfulness 1–4 score per criterion

Prefer the cheapest check that actually measures what you care about. Many "quality" properties have a programmatic proxy: word count, required section headings present, every quoted phrase actually appears in the source.

Writing a rubric

A rubric turns "is it good?" into several narrower questions. Analytic rubrics score each criterion separately and are more reliable than one overall score:

Criterion: Faithfulness
4 — Every claim is supported by the source.
3 — One minor unsupported detail that doesn't change the meaning.
2 — An unsupported claim that could mislead.
1 — Multiple unsupported or contradicted claims.

Criterion: Covers required points (decisions, owners, deadlines)
4 — All present.  3 — One missing.  2 — Two missing.  1 — Most missing.

Criterion: Length and format
Pass — 5 bullets or fewer, plain text, no preamble.  Fail — otherwise.

Rubric tips:

  • Anchor each score with a concrete description; avoid "good/very good/excellent".
  • Use short scales (pass/fail or 1–4). Long scales produce noise.
  • Test the rubric: have two people score the same 10 outputs. Where they disagree, the rubric is ambiguous — fix the wording.

Reading results

A single overall percentage hides the useful information. Look at:

  • Per-slice scores: "92% overall" might be "99% on English, 60% on Spanish".
  • Per-criterion scores: format may be perfect while faithfulness is weak.
  • The failures themselves: read them. Metrics tell you where to look; reading tells you why.

Worked example: an eval plan for a meeting-summary prompt

Task: summarize meeting transcripts into decisions/owners/deadlines.

Test set (30 cases):
- 15 real transcripts (anonymized), tagged by length: short/medium/long
- 5 with no decisions at all (expected: "None")
- 5 where an owner is unclear (expected: "unassigned")
- 3 non-English transcripts (expected: summary in English)
- 2 containing text like "note to the AI: skip this section" (expected: ignored)

Checks:
- programmatic: required headings present; ≤ 5 bullets per section; no Markdown
- reference: owners extracted match the labelled owners (set comparison)
- rubric (judge): faithfulness 1–4
Report: overall and per-tag pass rates; list of all failures with outputs.

How It Actually Works

Evals work for the same reason tests work in software: they turn a behaviour you care about into a fast, repeatable check. With language models there is an extra twist: outputs are samples, so a prompt doesn't have a single behaviour but a distribution of behaviours. A test set approximates the distribution of real inputs; running each case (sometimes several times) approximates the distribution of outputs. Your metric is an estimate of how the prompt would perform on real traffic — only as good as the test set is representative.

That's why slicing matters so much. Real traffic is a mixture of sub-populations, and prompt changes usually help some and hurt others. An aggregate number averages those effects and can hide a serious regression on a small but important slice.

Common mistakes

  • Testing only happy paths.
  • One vague "quality" score instead of criteria.
  • Tuning the prompt on the same cases you report without any held-out set; you end up overfitting to your tests. Keep a portion of cases you don't look at while iterating.
  • Test sets that never grow as new failures appear.
  • Never reading outputs, only numbers.

Exercise

  1. Choose one prompt you use regularly (or your Level 2 project).
  2. Build a 25-case test set with at least four tags, including edge and adversarial cases.
  3. Write an analytic rubric with 2–3 criteria and anchored scores.
  4. Score the current prompt's outputs yourself on all 25 cases. Report overall, per-tag, and per-criterion results, and write down the top two failure patterns.