Discussion about this post

User's avatar
Oli's avatar

Kinda hit this wall with a prompt pack I shipped — the judge model kept giving confident-but-wrong answers a pass. What fixed it was a small golden set of ~50 hand-labeled examples to calibrate the judge before trusting any score. The judge's only as good as the labels you feed it.

No posts

Ready for more?