PRODUCTION SCENARIO
A team uses Claude to generate support replies and wants to grade empathy on a 1-to-5 scale automatically. Their first grader prompt returns paragraphs of commentary that the pipeline cannot parse.
How should the grader be set up?
Answering here is anonymous. Nothing is saved unless you sign in.
Show answer and explanation
Answer: Use a different model as grader, define the scale, and ask for only the number
The guidance recommends using a different model to evaluate than the one that generated the output, defining the scale explicitly, and asking for only the number so results are parseable. Exact match does not fit a subjective quality such as tone.