What the data shows
am I too harsh on seventh graders' essays?no, it's the graders who are too generous- Good at order. The median rank correlation across tasks is 0.74. Sentence similarity (0.92), star ratings from reviews (0.88), job titles (0.84) and wine scores from tasting notes (0.84) are close to the labels.
- Rough on levels. Jev picks the exact grade 55% of the time on average, and within one grade most of the time (94% for sentence similarity, 96% for star ratings, less on AI replies and search). It's the fine lines between neighboring grades it misses.
- Weakest on AI replies and search. Ranking the helpfulness of AI assistant replies (0.60) and shopping search results (0.61) is harder: it hits the exact grade on only 38% of the AI replies.
- Harsh on essays. On seventh-grade essays, Jev orders them loosely (0.43) and grades them lower: 1.4 on average where the raters agreed on 1.9.
What it means, and what it doesn't
For ranking and sorting (which result goes first, which review is most positive), Jev's grades are useful as they are. For thresholds ("anything below a 3 goes to review"), calibrate on your own data first, because its lines between grades aren't the labels' lines.
The essay result is the one to act on: Jev reads student writing more harshly than the people who graded it, and not in quite the same order. "Jev rarely gives the top grade when reading others' judgments" looks at the same harshness from another angle.