Discussion about this post

User's avatar
Oli's avatar

Kinda hit this wall with a prompt pack I shipped — the judge model kept giving confident-but-wrong answers a pass. What fixed it was a small golden set of ~50 hand-labeled examples to calibrate the judge before trusting any score. The judge's only as good as the labels you feed it.

Mohamed F. Ahmed's avatar

Worth adding: the judge model itself needs versioning discipline. If you swap the underlying model powering your evaluator, your pass rates can shift even though your agent didn't change at all, and that's a brutal thing to debug six weeks later. Anyone here pin their judge model separately from their production model, or do you let both float?

2 more comments...

No posts

Ready for more?