Eval script lab
A lab to write a tiny offline eval runner: load fixtures → call model/system → score → print a report—and fail CI on regressions.
What it is
A lab to write a tiny offline eval runner: load fixtures → call model/system → score → print a report—and fail CI on regressions.
Why it matters
If it isn’t in a script, it isn’t a release gate.
How it works (plain)
JSONL fixtures → run → exact/rubric scores → compare to baseline file → exit non-zero on drop.
Try it
Encode 10 fixtures for one prompt; break the prompt on purpose; watch the script fail.
Myths
- ⚠️ Myth: Manual spot checks replace harnesses.
- ✓ Reality: Spot checks help; harnesses catch regressions.
Sources
- Course 18 evaluation; Course 08 prompt eval; Course 21 setup
