Evaluation of prompts
Treating prompts like code: version them, test them on a fixed set, and reject changes that regress quality.
What it is
Treating prompts like code: version them, test them on a fixed set, and reject changes that regress quality.
Why it matters
Prompt folklore doesn’t survive contact with production. Eval harnesses do.
How it works (plain)
Golden inputs → run prompt versions → score with rubrics/exact checks → keep winners → monitor online.
Try it
Create 10 fixtures for one prompt you care about; score two variants blind.
Myths
- ⚠️ Myth: A single clever prompt is done forever.
- ✓ Reality: Models and contexts change—re-test.
Sources
- Course 08 prompting; Course 18 evaluation
- Course 10 agent eval
