What the data shows
the short correct answerthe same answer, three times longer- Agreement: 83% overall. 85% on MT-Bench, 81% on HelpSteer2. When the judges were unanimous, 86%; when they were split, 73%, which is what you'd expect if Jev tracks the same qualities they do.
- Length: the same lean. When one answer is at least 1.5 times longer, Jev picks it 67% of the time; the judges, 70%. When the first answer is more than twice as long, Jev picks it 69% of the time and the judges 73%.
- Position: a small lean the other way. On HelpSteer2, Jev picks the first answer 44% of the time; the judges 51%.
What it means, and what it doesn't
As a judge of chatbot answers, Jev agrees with human judges about as often as you'd hope, and the famous "length bias" isn't an extra flaw here: it's inherited. People who judge AI answers prefer longer ones too, and Jev prefers them slightly less.
It doesn't mean length is a fair signal of quality; it means Jev won't correct for it. And these judges are experts and hired annotators; everyday users might weigh length differently.