Kinda hit this wall with a prompt pack I shipped — the judge model kept giving confident-but-wrong answers a pass. What fixed it was a small golden set of ~50 hand-labeled examples to calibrate the judge before trusting any score. The judge's only as good as the labels you feed it.
Worth adding: the judge model itself needs versioning discipline. If you swap the underlying model powering your evaluator, your pass rates can shift even though your agent didn't change at all, and that's a brutal thing to debug six weeks later. Anyone here pin their judge model separately from their production model, or do you let both float?
the golden dataset part hit close to home. built one for a support bot, passed everything, then a customer found a hallucinated refund policy on day two. our examples were all happy path — now we bake a few adversarial cases into every golden set.
The separation between a golden dataset, automated metrics, and judge-model review is a useful evaluation stack. One safeguard I’d add is slice-level reporting: break results out by task type, context length, and known failure mode so an improved aggregate score can’t hide a regression in the cases users care about most.
Kinda hit this wall with a prompt pack I shipped — the judge model kept giving confident-but-wrong answers a pass. What fixed it was a small golden set of ~50 hand-labeled examples to calibrate the judge before trusting any score. The judge's only as good as the labels you feed it.
Worth adding: the judge model itself needs versioning discipline. If you swap the underlying model powering your evaluator, your pass rates can shift even though your agent didn't change at all, and that's a brutal thing to debug six weeks later. Anyone here pin their judge model separately from their production model, or do you let both float?
the golden dataset part hit close to home. built one for a support bot, passed everything, then a customer found a hallucinated refund policy on day two. our examples were all happy path — now we bake a few adversarial cases into every golden set.
The separation between a golden dataset, automated metrics, and judge-model review is a useful evaluation stack. One safeguard I’d add is slice-level reporting: break results out by task type, context length, and known failure mode so an improved aggregate score can’t hide a regression in the cases users care about most.