4 Comments
User's avatar
Oli's avatar

Kinda hit this wall with a prompt pack I shipped — the judge model kept giving confident-but-wrong answers a pass. What fixed it was a small golden set of ~50 hand-labeled examples to calibrate the judge before trusting any score. The judge's only as good as the labels you feed it.

Mohamed F. Ahmed's avatar

Worth adding: the judge model itself needs versioning discipline. If you swap the underlying model powering your evaluator, your pass rates can shift even though your agent didn't change at all, and that's a brutal thing to debug six weeks later. Anyone here pin their judge model separately from their production model, or do you let both float?

Oli's avatar

the golden dataset part hit close to home. built one for a support bot, passed everything, then a customer found a hallucinated refund policy on day two. our examples were all happy path — now we bake a few adversarial cases into every golden set.

Promptslove's avatar

The separation between a golden dataset, automated metrics, and judge-model review is a useful evaluation stack. One safeguard I’d add is slice-level reporting: break results out by task type, context length, and known failure mode so an improved aggregate score can’t hide a regression in the cases users care about most.