What the data shows
Jev: 70%volunteers: 71%this reply fails the task- Personal attacks: close to a panel. Where about 80% of raters saw an attack, Jev's average is 86%; where none did, 3%. Its rank correlation with the rater share is 0.78.
- Chatbot replies: right in the middle. Where about 71% of volunteers said the reply failed, Jev gives 70%.
- The ends are where it drifts. For replies no volunteer faulted, Jev still gives 22%. For comments every hate speech rater flagged, it gives only 70%.
- Off by 0.05 to 0.12 on average, depending on the dataset.
What it means, and what it doesn't
Where most raters said yes, Jev's probabilities can be read the way you'd read a small panel's vote: 70% means "debatable, leaning yes". That's useful, because it means its uncertainty carries information about human disagreement. Lower down the scale it runs higher than the panel.
It's less trustworthy at the ends. A reply that every volunteer accepted still gets a one-in-five "fails" from Jev, and unanimous hate speech gets only 70%: on those, its number is more hesitant than the crowd. And these are three datasets on moderation and helpfulness; the same calibration isn't guaranteed on other kinds of judgments.