We grade support replies with a model. Early on, the grader passed replies that quoted the wrong refund window, because the window sounded plausible.
Now every grader has a small labeled set of its own: 60 replies a person marked pass or fail. When we change the grader prompt, we run it on that set first.
Agreement with people is the only grader score we trust.
