Why a clean average can hide a failed AI answer
A batch average does not tell a reviewer which answer must not ship. Separate sample severity from an overall score.
Consider two fictional outputs. One has harmless extra punctuation. The other gives a 30-day refund window when the supplied policy says 14 days. An average can make both look like small deductions, but the second answer changes what the user can do.
A compact review record has six parts: the request, supplied evidence, model output, severity, reason, and suggested fix. Keep critical and major failures visible in their own counts. Do not treat a handful of teaching examples as a statistically useful test set.
Before a real review, ask two people to label the same examples independently. Discuss disagreement, improve the definitions, version the rubric, and re-review affected examples. Agreement means reviewers used the labels consistently; it does not prove both were correct.
I am sharing this on behalf of JudgeMyAI, a human-led LLM evaluation service. Our evaluation guide explains the surrounding workflow: https://judgemyai.com/evaluation-guide/ . AI assistance was used in preparing this post.
Comments
No comments yet
Be the first to share your thoughts